Data decompression method, data compression method and convolution operation device

By using data decompression and compression methods in the convolution operation module, the first-level compression algorithm is used to process the input data block, which solves the problem of insufficient cache space and improves the efficiency and speed of convolution operation.

CN111884658BActive Publication Date: 2025-05-20GLENFLY TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010656506.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-09
Publication Date
2025-05-20
Estimated Expiration
2040-07-09

AI Technical Summary

Technical Problem

The cache space of the convolutional computing module is limited, which makes it impossible to cache all required computing data, affecting the computing speed.

Method used

By using data decompression method and data compression method in the convolution operation module, the input data block is compressed and decompressed using a primary compression algorithm to reduce the storage space requirement and thus cache more operation data.

Benefits of technology

The number of times the convolution operation module is paused is reduced, and the computing efficiency and speed is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111884658B_ABST
    Figure CN111884658B_ABST
Patent Text Reader

Abstract

The present invention provides a data decompression method, a data compression method and a convolution operation device. One embodiment of the present invention proposes a data decompression method, which is used in a convolution operation device to decompress an input data block, and the data decompression method comprises: reading the input data block; and performing a first-level decompression on the input data block; wherein the first-level decompression is used to decompress the input data block whose data format is a first-level compression algorithm format; wherein the first-level compression algorithm format comprises: a mask field, which is used to identify the position of non-zero elements in the input data block and the number of elements in the input data block.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a data decompression method, a data compression method and a convolution operation device, and particularly relates to a data decompression method, a data compression method and a convolution operation device for decompressing / compressing an input data block. Background Art

[0002] Convolutional Neural Networks (CNN) are currently the main force in the development of the field of deep neural networks and are very accurate in image recognition. A typical convolutional neural network includes many layers of operations, such as a convolution layer, an activation layer, a pooling layer, and a fully connected layer.

[0003] Using a convolution operation module (a hardware module, such as a CNN accelerator, etc.) independent of the CPU (Central Processing Unit) can effectively improve the speed of convolution operations. However, the cache space in the convolution operation module for caching operation data (including input data and convolution kernels, etc.) is limited. When performing convolution operations, not all the operation data used in the current convolution layer can be cached in the convolution operation module. Therefore, if the operation data required for convolution operations has not been cached in the convolution operation module, the convolution operation module will pause the convolution operation and load the required operation data from the memory outside the convolution operation module. Only after loading the required operation data can the convolution operation continue, thus affecting the operation speed of the convolution operation module.

[0004] Therefore, how to cache more operation data when the cache space of the convolution operation module is limited, and how to make each loaded operation data more, so as to reduce the number of times the convolution operation module pauses, thereby improving the operation efficiency of the convolution operation module, has become one of the problems to be solved in this field. Summary of the Invention

[0005] In view of this, the present invention provides a data decompression method, a data compression method and a convolution operation device, which can cache more operation data in the convolution operation module to reduce the number of times the convolution operation module pauses, thereby improving the operation efficiency of the convolution operation module.

[0006] An embodiment of the present invention provides a data decompression method for a convolution operation device to decompress an input data block. The data decompression method includes: reading the input data block; and performing primary decompression on the input data block. Wherein, the primary decompression is used to decompress the input data block with a data format of a primary compression algorithm format. Wherein, the primary compression algorithm format includes: a mask field for identifying the positions of non-zero elements in the input data block and the number of elements in the input data block.

[0007] An embodiment of the present invention provides a data compression method for a convolution operation device to compress an input data block. The data compression method includes: generating the input data block; and performing primary compression on the input data block. Wherein, the primary compression compresses the input data block into data with a data format of a primary compression algorithm format. Wherein, the primary compression algorithm format includes: a mask field for identifying the positions of non-zero elements in the input data block and the number of elements in the input data block.

[0008] An embodiment of the present invention provides a convolution operation device for decompressing an input data block. The convolution operation device includes: a cache for storing the input data block; and a primary processing module for reading the input data block from the cache and performing primary decompression on the input data block. Wherein, the primary decompression is used to decompress the input data block with a data format of a primary compression algorithm format. Wherein, the primary compression algorithm format includes: a mask field for identifying the positions of non-zero elements in the input data block and the number of elements in the input data block.

[0009] An embodiment of the present invention provides a convolution operation device for compressing an input data block. The convolution operation device includes: a cache for storing the input data block; and a data processing module for performing primary compression on the input data block and storing the input data block after primary compression into the cache. Wherein, the primary compression compresses the input data block into data with a data format of a primary compression algorithm format. Wherein, the primary compression algorithm format includes: a mask field for identifying the positions of non-zero elements in the input data block and the number of elements in the input data block.

[0010] By means of the data decompression method, data compression method and convolution operation device of the present application, more input data blocks can be cached in the convolution operation device by compressing and storing the input data blocks, thereby reducing the number of pauses of the convolution operation module, and thus improving the operation efficiency of the convolution operation module. Description of the Drawings

[0011] Figure 1Schematic diagram of a convolutional neural network 100 according to an embodiment of the present invention.

[0012] Figure 2 Schematic diagram of the convolution operation of the Nth convolutional layer and the (N + 1)th convolutional layer in the convolutional neural network 100 according to an embodiment of the present invention.

[0013] Figure 3A Schematic diagram of the block convolution operation when the convolution kernel is 1*1 according to an embodiment of the present invention.

[0014] Figure 3B Schematic diagram of the overlap of the input data block in the up and down directions when the convolution kernel is 3*3 during the convolution operation according to an embodiment of the present invention.

[0015] Figure 3C Schematic diagram of the overlap of the input data block in the left and right directions when the convolution kernel is 3*3 during the convolution operation according to an embodiment of the present invention.

[0016] Figure 3D Schematic diagram of the overlap of the input data block in the upper left - lower right direction when the convolution kernel is 3*3 during the convolution operation according to an embodiment of the present invention.

[0017] Figure 3E Schematic diagram of the overlap of the input data block in the lower left - upper right direction when the convolution kernel is 3*3 during the convolution operation according to an embodiment of the present invention.

[0018] Figure 4 Schematic diagram of the feature map block when the convolution kernel is k*k and the convolution stride is s during the convolution operation according to an embodiment of the present invention.

[0019] Figure 5 Block diagram of a computing device 500 including a convolution operation module according to an embodiment of the present invention.

[0020] Figure 6A Schematic diagram of the data stored in the memory 520 of the computing device 500 according to an embodiment of the present invention.

[0021] Figure 6B More detailed block diagram of the computing device 500 according to an embodiment of the present invention.

[0022] Figure 6C Processing flow of compressing the input feature map of the Nth convolutional layer in two - levels and writing it into the memory according to an embodiment of the present invention.

[0023] Figure 6D Processing flow of the computing device 500 generating an output feature map according to an embodiment of the present invention.

[0024] Figure 6E The processing flow for generating an output feature map for the computing device 500 illustrated in another embodiment of the present invention.

[0025] Figure 6F-1 to 6F-2 A more detailed processing flow for the computing device 500 illustrated in an embodiment of the present invention to generate an output feature map in a left-to-right and top-to-bottom order.

[0026] Figure 7 The processing flow for the computing device 500 illustrated in an embodiment of the present invention to decompress an input data block.

[0027] Figure 8 The block diagram of the computing device 800 including a convolution operation module illustrated in another embodiment of the present invention.

[0028] Figure 9A The schematic diagram of the data stored in the memory 820 of the computing device 800 illustrated in an embodiment of the present invention.

[0029] Figure 9B A more detailed block diagram of the computing device 800 illustrated in an embodiment of the present invention.

[0030] Figure 9C The processing flow for writing the input feature map of the Nth convolution layer into the cache after first-level compression illustrated in an embodiment of the present invention.

[0031] Figure 9D The processing flow for the computing device 800 to generate an output feature map illustrated in an embodiment of the present invention.

[0032] Figure 9E The processing flow for the computing device 800 to generate an output feature map illustrated in another embodiment of the present invention.

[0033] Figure 9F-1 to 9F-2 A more detailed processing flow for the computing device 800 to generate an output feature map illustrated in an embodiment of the present invention.

[0034]

Symbol Explanation

[0035] 100 Convolutional neural network

[0036] 110 Input data

[0037] 120 Feature extraction stage

[0038] 121 - 12X Convolutional layer

[0039] 130 Classification stage

[0040] 131-13Y Fully Connected Layer

[0041] 140 Output Data

[0042] 210, 230, 250 Feature Map Sets

[0043] 220, 240 Weights

[0044] 221, 223, 241, 243, 245 Convolution Kernel Groups

[0045] 2211-2213, 2231-2233 Convolution Kernels

[0046] 310A-310E, 410 Input Feature Maps

[0047] 313A-313E, 413 Convolution Kernels

[0048] 315A-315E, 415 Output Feature Maps

[0049] 1-10 Number of Columns or Rows

[0050] W, w1-w3 Width

[0051] H, h1-h3 Height

[0052] k Side Length of Convolution Kernel

[0053] s Convolution Stride

[0054] 500, 800 Computing Devices

[0055] 520, 820 Memories

[0056] 530, 830 Convolution Operation Modules

[0057] 531 Configuration Register

[0058] 538 Secondary Processing Module

[0059] 539, 839 Data Processing Modules

[0060] 534 Primary Processing Module

[0061] 535 Segmentation Module

[0062] 537, 837 Compression Modules

[0063] 532, 832 Caches

[0064] 5321, 5323, 8322 Cache Segments

[0065] 536 Arithmetic Unit

[0066] 5361 - 536Z Arithmetic Unit

[0067] 521, 523, 525, 527, 821, 823 Storage Segments

[0068] M, N Numbers

[0069] 52111 - 521M1 Main Area

[0070] 52112 - 521M2 Sub - Area

[0071] 53211 - 5321M Input Feature Map Cache Segment

[0072] 532111 - 5321M1, 832111 - 8321M1 Main Cache Segments

[0073] 532113 - 5321M3, 832113 - 8321M3 Sub - Cache Segments

[0074] 5342 Scratchpad

[0075] 53421 Main Scratchpad Segment

[0076] 53423 Sub - Scratchpad Segment

[0077] 53425 Convolution Kernel Group Scratchpad Segment

[0078] 534211 - 53421M Main Area

[0079] 5342311 - 534231M, 5342331 - 534233M Sub - Areas

[0080] 534251 - 53425M Convolution Kernel Group

[0081] Steps S601C, S603C, S605C, S607C

[0082] Steps S603D, S605D, S607D, S609D

[0083] Steps S601E, S603E, S605E, S607E, S609E

[0084] Steps S601F, S603F, S605F, S607F, S609F, S613F, S615F, S617F, S619F, S621F, S623F, S625F, S627F, S629F

[0085] Steps S701, S703

[0086] Steps S901C, S903C, S907C

[0087] Steps S903D, S905D, S907D, S909D

[0088] Steps S905E, S907E, S909E

[0089] Steps S901F, S913F, S915F, S917F, S919F, S921F, S923F, S925F, S927F, S929F Detailed implementation manners

[0090] The following describes the implementation manners for realizing the present invention, which aims to describe the basic spirit of the present invention, but is not used to limit the present invention. The actual content of the invention must refer to the scope of the following claims.

[0091] It must be understood that the words such as "comprising" and "including" used in this specification are used to indicate the existence of specific technical features, numerical values, method steps, operations, elements and / or components, but do not exclude the addition of more technical features, numerical values, method steps, operations, elements, components, or any combination of the above.

[0092] In the claims, words such as "first", "second", and "third" are used to modify the elements in the claims, and are not used to indicate a priority order, precedence relationship, or that one element precedes another element, or the time sequence when performing method steps, but are only used to distinguish elements with the same name.

[0093] Two lossless compression algorithms are used in the technical solutions in the present disclosure, namely primary compression and secondary compression. For the convenience of subsequent description, we first describe these two compression algorithms. The secondary compression algorithm can be the Huffman algorithm, the LZW (Lenpel-Ziv & Welch) algorithm, etc. Correspondingly, the secondary compression algorithm format is the format of algorithms such as the Huffman algorithm and the LZW (Lenpel-Ziv & Welch) algorithm. In the present disclosure, we generally use the secondary compression algorithm to perform another compression on the data after primary compression to further improve the compression ratio.

[0094] The primary compression algorithm can be used to compress a matrix containing a relatively large number of elements with a value of 0. The primary compression algorithm format is as follows (including three fields, where the "+" indicates that the two fields before and after are closely connected without other data in between):

[0095] [Length]+[Mask]+[DesData]

[0096] The DesData field represents the target data field. The DesData field contains all the elements in the matrix whose values are not 0. The order of all elements in the DesData field is the same as their order in the matrix (there are two ways to arrange the order of elements in a two-dimensional matrix: 1. In the order from left to right and top to bottom; 2. In the order from top to bottom and left to right).

[0097] The Mask field represents the mask field, and the length of the Mask field can be set according to the number of elements in the compressed matrix. The Mask field has two functions. The first function is to represent the number of elements in the compressed matrix; the second function is to mark the positions of non-zero elements in the compressed matrix. There are two ways to use the Mask field to represent the number of elements in the compressed matrix. The first way is to set the length of the Mask field to be equal to the number of elements in the compressed matrix (the situation of using the first way will be described later); the second way is to set the length of the Mask field to be greater than the number of elements in the compressed matrix, and set the value of the bit in the Mask field corresponding to the last element in the compressed matrix to 1, and set the values of the bits in the Mask field that have no corresponding relationship with the elements in the compressed matrix to 0. In this way, the number of elements in the compressed matrix can be calculated according to the position of the last bit with a value of 1 in the Mask field (the situation of using the second way will be described later). In the present disclosure, many matrices need to be compressed. When the number of elements in all the compressed matrices is the same, the length of the Mask field (the length of the Mask field is the number of bits contained in the Mask field, the same below) can be set to the number of elements in the compressed matrix. For example, when the width and height of all the compressed matrices are m and n respectively (that is, the compressed matrix contains m columns and n rows of elements, and m and n can be the same or different integers greater than 0), set the length of the Mask field to m * n (* represents the multiplication sign, the same below) bits. Each element in the compressed matrix corresponds to each bit in the Mask field. Each bit with a value of 0 in the Mask field corresponds to an element with a value of 0 in the compressed matrix, and each bit with a value of 1 in the Mask field corresponds to an element with a value of not 0 in the compressed matrix. When the value of an element in the compressed matrix is not 0, the value of this element will be stored in the corresponding position in the DesData field, and the value of the corresponding bit in the Mask field is 1. It should be noted that in another embodiment, the bits with a value of 0 in the Mask field correspond to the elements with a value of not 0 in the compressed matrix, and the bits with a value of 1 correspond to the elements with a value of 0 in the compressed matrix.

[0098] The Length field represents the length field, which is used to indicate the length of the DesData field (the length of the DesData field refers to the number of elements in the DesData field, the same below). There are two ways to use the Length field to represent the length of the DesData field, which are respectively called the first length representation method and the second length representation method. In the first length representation method, the value of the Length field is equal to the length of the DesData field, and the maximum length value that the Length field can indicate is equal to the maximum value of the Length field. For example, for a Length of 1 byte, Length can represent the length of the DesData field in the range of 0 - 255. In the first length representation method, when the length of the Length field is 1 byte, if the length of the DesData field exceeds 255 (such as 260), it cannot be represented by the Length field. If you want to represent a length greater than 255, you need to use a Length field with a larger length (for example, changing the length of the Length field to 2 bytes can represent a length of 260), but this will increase the storage space occupied by the Length field. To solve this problem, the second length representation method of using the Length field to represent the length of the DesData field is proposed in this disclosure. In the second length representation method, each value of the Length field indicates a specific length value, and the maximum number of elements that the Length field can indicate is greater than the maximum value of the Length field. For example, a Length field with a length of 2 bits can represent 4 length values, and the length value represented by each value of the Length field can be preset according to actual needs. For example, in an embodiment, the value of the Length field is

[00] 2 ( 2 It means that the numbers in [] are binary numbers, the same below) represents that the length of the DesData field is 8, and the value of the Length field is

[01] 2 represents that the length of the DesData field is 12, and the value of the Length field is

[10] 2 represents that the length of the DesData field is 18, and the value of the Length field is

[11] 2The length of the DesData field in the time representation is 24. If the number of non-zero elements in the compressed matrix is different from the length that can be represented by the value of the Length field (i.e., the number of non-zero elements in the compressed matrix is not one of 8, 12, 18, or 24), then the value of the Length field corresponding to the smallest length value that is greater than the number of non-zero elements in the compressed matrix and can be represented by the value of the Length field can be selected. For example, when the number of non-zero elements in the compressed matrix is 6, the smallest length greater than 6 that can be represented by the value of the Length field is 8 (the corresponding value of the Length field is

[00] 2 ), so the value of the Length field

[00] 2 is selected. Since the value of the Length field

[00] 2 indicates that the length of the DesData field is 8, when performing compression processing, the DesData field will contain 8 elements, of which the first 6 elements are the non-zero elements in the compressed matrix, and the last 2 elements can be set to 0 or other values; 6 bits corresponding to the 6 elements in the compressed matrix in the Mask field are set to 1, and the other bits are set to 0. When performing decompression processing, the compressed matrix can be generated according to the positions of the bits with the value of 1 in the Mask field and the element values in the DesData field corresponding to the bits with the value of 1 in the Mask field.

[0099] For ease of understanding, the following example illustrates how to compress a matrix using the first-level compression algorithm. Assume that the compressed matrix Matrix1 is as follows (assuming the width (i.e., m) of the matrix is 5 and the height (i.e., n) is 4):

[0100] 0 0 8 0 0 0 0 0 0 5 0 0 9 10 0 0 0 0 4 0

[0101] When compressing the compressed matrix Matrix1 using the first length representation method with the value of the Length field representing the length of the DesData field, set the length of the Length field to 1 byte and the length of the Mask to 20 bits (since the compressed matrix Matrix1 has 20 (5 * 4 = 20) elements, the length of the Mask is set to 20 bits). The data after the first-level compression of the compressed matrix Matrix1 is (compressed in the order from left to right and top to bottom in the matrix):

[0102] [5] 10 +[00100, 00001, 00110, 00010] 2 +[8, 5, 9, 10, 4] 10

[0103] Among them, [] 10It means the number in [] is a decimal number, 2 It means the number in [] is a binary number. [5] 10 The 5 in [] means that the DesData field contains 5 elements.

[0104] Suppose each element in the compressed matrix Matrix1 occupies 1 byte of storage space. Before compression, Matrix1 needs to occupy 20 bytes of storage space. After the first-level compression, the Length field occupies 1 byte of storage space, the Mask field occupies 3 bytes (20 bits) of storage space, and the DesData field occupies 5 bytes of storage space. That is, after the first-level compression, Matrix1 needs to occupy a total of 9 bytes of storage space. Therefore, in this example, when using the first length representation method, the compression ratio is 9 / 20.

[0105] When using the second length representation method to compress the compressed matrix Matrix1 with the value of the Length field representing the length of the DesData field, set the length of the Length field to 2 bits and the length of the Mask field to 20 bits. When the value of the Length field is

[00] 2 it represents that the length of the DesData field is 8, and the value of the Length field is

[01] 2 it represents that the length of the DesData field is 12, and the value of the Length field is

[10] 2 it represents that the length of the DesData field is 18, and the value of the Length field is

[11] 2 it represents that the length of the DesData field is 24. The data of the compressed matrix Matrix1 after the first-level compression is (compressed in the order from left to right and from top to bottom in the matrix):

[0106]

[00] 2 +[00100, 00001, 00110, 00010] 2 +[8, 5, 9, 10, 4, 0, 0, 0] 10

[0107] Among them, 10 It means the number in [] is a decimal number, 2 It means the number in [] is a binary number.

[00] 2 It means that the DesData field contains 8 elements, [00100, 00001, 00110, 00010] 2 There are only 5 1s in [], indicating that the compressed matrix Matrix1 only contains 5 elements with non-zero values. When performing decompression processing, [8, 5, 9, 10, 4, 0, 0, 0]10 The last 3 elements will be ignored.

[0108] Assume that each element in the compressed matrix Matrix1 occupies 1 byte of storage space. Before compression, Matrix1 requires 20 bytes of storage space. After the first-level compression, the Length field occupies 2 bits of storage space, and the Mask field occupies 20 bits of storage space, that is, the Length field and the Mask field together occupy 3 bytes of storage space (22 bits in total); the DesData field occupies 8 bytes of storage space. That is, after the first-level compression, Matrix1 requires a total of 11 bytes of storage space. Therefore, in this example, when using the second length representation method, the compression ratio is 11 / 20.

[0109] In another embodiment, when the number of elements in multiple compressed matrices is different (i.e., some compressed matrices have more elements and some have fewer elements), in order to simplify the decompression process, the length of the Mask field can be set to the number of elements in the compressed matrix with the most elements. In this embodiment, since the length of the Mask field is no longer the same as the number of elements in the compressed matrix, we can no longer use the length of the Mask field to represent the number of elements in the compressed matrix, and a new mechanism for representing the number of elements in the compressed matrix is needed. For this purpose, we use the bit in the Mask field corresponding to the last element in the compressed matrix as a marker for calculating the number of elements in the compressed matrix (set the value of this bit to 1). Specifically, when performing matrix compression processing, regardless of whether the last element in the compressed matrix is 0 or not, the bit in the Mask field corresponding to it is set to 1, and all bits after this bit in the Mask field are set to 0. Therefore, by subtracting the number of bits after the last bit with a value of 1 in the Mask field from the total number of bits in the Mask field, the number of elements in the compressed matrix can be obtained. For elements in the compressed matrix other than the last element, if the value is 0, the corresponding bit in the Mask field is set to 0; if the value is not 0, the corresponding bit in the Mask field is set to 1. In this way, when performing matrix decompression processing, the number of elements in the compressed matrix can be obtained according to the position of the last bit with a value of 1 in the Mask field. For example, when the size of the compressed matrix with the most elements is 6*4 (i.e., it contains 24 elements), the length of the Mask field is set to 24 bits (bit). Each element in the compressed matrix corresponds to a bit in the Mask field. Each element with a value of 0 in the compressed matrix other than the last element corresponds to a bit with a value of 0 in the Mask field, each element with a value not equal to 0 in the compressed matrix other than the last element corresponds to a bit with a value of 1 in the Mask field, and the last element in the compressed matrix (with a value of 0 or not 0) corresponds to the last bit with a value of 1 in the Mask field. In this embodiment, since the bit in the Mask field corresponding to the last element of the compressed matrix must be 1, when performing decompression processing, it is impossible to determine whether the last element of the compressed matrix is not 0 based on the value of this bit. Therefore, we need to store the value of the last element of the compressed matrix in the DesData field (even if its value is 0).

[0110] In this embodiment, when compressing the compressed matrix Matrix1 using the first length representation method, where the length of the DesData field is represented by the value of the Length field, first set the length of the Length field to 1 byte and the length of the Mask to 24 bits (since the compressed matrix with the most elements among multiple compressed matrices contains 24 elements, the length of the Mask is set to 24 bits). The data of the compressed matrix Matrix1 after the first-level compression is (compressed in the order from left to right and top to bottom in the matrix):

[0111] [6] 10 +[00100, 00001, 00110, 00011, 0000] 2 +[8, 5, 9, 10, 4, 0] 10

[0112] Among them, 10 it means the numbers in [] are decimal numbers, 2 and it means the numbers in [] are binary numbers. [6] 10 The 6 in [] indicates that the DesData field contains 6 elements. The last element 0 in the DesData field is the last element in the compressed matrix Matrix1, and the corresponding bit in the Mask field is the last bit with a value of 1 (i.e., the 20th bit in the Mask field). The last bit with a value of 1 in the Mask field is the 20th bit in the Mask field, indicating that the compressed matrix Matrix1 contains 20 elements.

[0113] Assume that each element in the compressed matrix Matrix1 occupies 1 byte of storage space. Before compression, Matrix1 needs to occupy 20 bytes of storage space. After the first-level compression, the Length field occupies 1 byte of storage space, the Mask field occupies 3 bytes (24 bits) of storage space, and the DesData field occupies 6 bytes of storage space. That is, after the first-level compression, Matrix1 altogether needs to occupy 10 bytes of storage space. Therefore, in this example, the compression ratio is 10 / 20, that is, the compression ratio is 1 / 2.

[0114] Now please refer to Figure 1 , Figure 1 which is a schematic diagram of the convolutional neural network 100 illustrated according to an embodiment of the present invention. As Figure 1As shown, the convolutional neural network 100 includes a feature extraction stage 120 and a classification stage 130, and the input data 110 comes from outside the neural network 100. Taking an RGB image as an example, the input data 110 includes 3 images: the image of the R channel of the RGB image, the image of the G channel, and the image of the B channel. Taking a grayscale image as an example, the input data 110 only includes 1 image.

[0115] The feature extraction stage 120 includes at least one convolutional layer for extracting features from the input data 110. The input data 110 is the input data of the first convolutional layer 121 of the feature extraction stage 120. After the first convolutional layer 121 performs a convolution operation (i.e., a feature extraction operation) on the input data, it generates the output data of the first convolutional layer 121. The output data of the first convolutional layer 121 can be used as the input data of the second convolutional layer 122 (i.e., the next convolutional layer). After the second convolutional layer 122 performs a convolution operation (i.e., a feature extraction operation) on the input data, it generates the output data of the second convolutional layer 122 (i.e., the input data of the next convolutional layer). And so on, the Xth convolutional layer 12X performs a convolution operation on the input data from the previous convolutional layer and generates the output data of the Xth convolutional layer 12X. The output data of the Xth convolutional layer 12X is sent to the classification stage 130 for classification processing.

[0116] In a neural network, there is often an activation layer (not shown) after many convolutional layers. The activation layer performs activation processing on the output data of the convolutional layer and then sends it to the next convolutional layer for convolution operation. After activation processing, a large amount of sparse data will appear in the neural network (i.e., the data contains a large number of elements with a value of 0). Under the first-level compression method disclosed in the present invention, since only non-zero elements are stored, the data storage space required for performing convolution operations can be greatly reduced. Moreover, the data appearing in the neural network includes input feature maps, output feature maps, convolution kernels, etc. Among them, the input feature maps, regions of the input feature maps, output feature maps, and regions of the output feature maps, etc. all belong to the matrices mentioned above and can all be compressed using the first-level compression algorithm and the second-level compression algorithm. Before storing the large amount of sparse data appearing in the neural network, by using the first-level compression algorithm proposed in the present disclosure to compress it, a large amount of storage space can be saved and the data transmission efficiency can be improved.

[0117] In another embodiment, there is a pooling layer after some convolutional layers (or activation layers). The pooling layer performs pooling processing on the output data of the convolutional layer (or activation layer) and then sends it to the next convolutional layer for convolution operation.

[0118] The output data of the feature extraction stage 120 is sent to the classification stage 130 as the input data for processing in the classification stage 130. The classification stage 130 includes a plurality of fully connected layers (the first fully connected layer 131 to the Y-th fully connected layer 13Y). After receiving the input data (i.e., the output data of the feature extraction stage 120), the first fully connected layer 131 to the Y-th fully connected layer 13Y sequentially process the received input data, and finally generate the output data 140. The output data 140 is the data output from the neural network 100 to the outside.

[0119] After the image in the input data 110 undergoes the convolution operation (i.e., the feature extraction operation) of the first convolutional layer in the feature extraction stage 120, the generated image is called a feature map. In the feature extraction stage 120, starting from the second convolutional layer until the X-th convolutional layer, the image included in the input data of each convolutional layer is called an input feature map, and the image included in the output data of each convolutional layer is called an output feature map. For the convenience of description, in the present disclosure, the image in the input data 110 is also referred to as an input feature map.

[0120] Figure 2 FIG. is a schematic diagram of the convolution operations of the N-th convolutional layer and the N + 1-th convolutional layer in the convolutional neural network 100 according to an embodiment of the present invention. As Figure 2 shown, the feature map set 210 is the input data of the N-th convolutional layer of the convolutional neural network 100, and the feature map set 230 is the output data of the N-th convolutional layer of the convolutional neural network 100. The feature map set 230 is also the input data of the N + 1-th convolutional layer of the convolutional neural network 100, and the feature map set 250 is the output data of the N + 1-th convolutional layer of the convolutional neural network 100. The convolutional kernel group set 220 is the set of convolutional kernel groups of the N-th convolutional layer of the convolutional neural network 100, and the convolutional kernel group set 240 is the set of convolutional kernel groups of the N + 1-th convolutional layer of the convolutional neural network 100.

[0121] The set of feature maps 210 includes feature maps 211, 213, and 215. The set of feature maps 230 includes feature maps 231 and 233, and the set of convolution kernel groups 220 includes convolution kernel groups 221 and 223. The convolution kernel group 221 includes convolution kernels 2211, 2212, and 2213. In the convolution operation of the Nth convolutional layer, each convolution kernel in the convolution kernel group 221 is respectively subjected to a convolution operation with the corresponding feature map in the set of feature maps 210 to generate the feature map 231 in the set of feature maps 230. Specifically, the feature map 211 is subjected to a convolution operation with the convolution kernel 2211 to generate a first feature map (not shown), the feature map 213 is subjected to a convolution operation with the convolution kernel 2212 to generate a second feature map (not shown), the feature map 215 is subjected to a convolution operation with the convolution kernel 2213 to generate a third feature map (not shown), and then the values of the pixels at the same positions in the first feature map, the second feature map, and the third feature map are added to generate the pixel value at the corresponding position in the feature map 231 (for example, adding the value of the pixel at the first row and the first column of the first feature map, the value of the pixel at the first row and the first column of the second feature map, and the value of the pixel at the first row and the first column of the third feature map to generate the pixel value at the first row and the first column of the feature map 231; and so on, all the pixel values in the feature map 231 can be generated). Similarly, the convolution kernels 2231, 2232, and 2233 in the convolution kernel group 223 are respectively subjected to a convolution operation with the corresponding feature maps 211, 213, and 215 in the set of feature maps 210, and then the feature map 233 in the set of feature maps 230 is generated according to the results of the convolution operation. According to the requirements of actual applications, a pooling layer (not shown) can be added between the Nth convolutional layer and the (N + 1)th convolutional layer, and the generated feature maps 231 and 233 are subjected to pooling processing and then output, and then the (N + 1)th convolutional layer performs a convolution operation on the feature maps 231 and 233 after pooling processing.

[0122] Similar to the convolution operation of the Nth convolutional layer, in the convolution operation of the (N + 1)th layer, the convolution kernel groups 241, 243, and 245 in the set of convolution kernel groups 240 are respectively subjected to a convolution operation with the feature maps 231 and 233 in the set of feature maps 230 to generate the feature maps 251, 253, and 255 in the set of feature maps 250.

[0123] From Figure 2 it can be seen that the number of input feature maps in each convolutional layer is the same as the number of convolution kernels in the convolution kernel group, and each convolution kernel group corresponds to an output feature map. All the input feature maps are required to calculate each output feature map. Taking the Nth convolutional layer as an example, when calculating the output feature map 231, all the convolution kernels in the convolution kernel group 221 and all the input feature maps 211, 213, and 215 in the set of feature maps 210 are required.

[0124] Since the width and height of the input data blocks that the convolution operation device can process in parallel are fixed (for example: 5*4), when performing convolution operations using the convolution operation device, if the width or height of the input feature map is greater than the width or height of the input data blocks that the convolution operation device can process in parallel, the input feature map needs to be first divided into multiple input data blocks. Then, the input data blocks are sent to the convolution operation device for convolution operations to generate output data blocks, and finally, the generated output data blocks are concatenated in order to form an output feature map. The following will be combined with Figure 3A-3E to analyze various situations when dividing the input feature map into input data blocks (assuming that there is only 1 input feature map, 1 convolution kernel, and 1 output feature map in the convolution layer where the convolution operation is performed in the example of Figure 3A-3E ). In the following analysis process, it is assumed that the width and height of the input data blocks that the convolution operation device can process in parallel are 5*4, and the convolution stride is assumed to be 1.

[0125] Now, please refer to Figure 3A , Figure 3A which is a schematic diagram of block convolution operation when the convolution kernel is 1*1 according to an embodiment of the present invention. As Figure 3A shown, 310A is the input feature map, 313A is the convolution kernel, and 315A is the output feature map generated after the convolution operation of the input feature map 310A and the convolution kernel 313A. Each square in the input feature map 310A and the output feature map 315A represents a feature value (i.e., pixel value), and each square in the convolution kernel 313A represents a weight value. The size of the input feature map 310A is 10*8. Since the size of the convolution kernel is 1*1, each feature value in the output feature map 315A is the product obtained by multiplying the feature value at the same coordinate in the input feature map 310A by the weight value in the convolution kernel 313A. Therefore, each feature value in the output feature map 315A corresponds one-to-one with each feature value in the input feature map 310A, that is, the size of the output feature map 315A is the same as that of the input feature map 310A, both being 10*8.

[0126] As Figure 3A shown, when the convolution kernel is 1*1, in order to generate the output data block with a right diagonal line (i.e., " / ", the same below) in the output feature map 315A, it is necessary to perform a convolution operation on the input data block with a right diagonal line in the input feature map 310A and the convolution kernel 313A. In order to generate the output data block with a left diagonal line (i.e., "\", the same below) in the output feature map 315A, it is necessary to perform a convolution operation on the input data block with a left diagonal line in the input feature map 310A and the convolution kernel 313A. Therefore, when the convolution kernel is 1*1, the two input data blocks in the input feature map 310A required to generate two adjacent and non-overlapping output data blocks in the output feature map 315A are also adjacent and non-overlapping.

[0127] Now please refer to Figure 3B , Figure 3B which is a schematic diagram showing the overlapping situation of the input data block in the up-down direction when the convolution kernel is 3*3 during convolution operation according to an embodiment of the present invention. As Figure 3B shown, 310B is the input feature map, 313B is the convolution kernel, and 315B is the output feature map generated after the convolution operation of the input feature map 310B and the convolution kernel 313B. Different from Figure 3A , Figure 3B the size of the convolution kernel 313B used during convolution operation is 3*3. As Figure 3B shown, when the convolution kernel is 3*3, the number of rows of the output feature map 315B is 2 rows less than that of the input feature map 310B, and the number of columns is 2 columns less (the size of the output feature map 315B is 8*6, and the size of the input feature map 310B is 10*8). The convolution operation process for generating the output feature map 315B is as follows: Starting from the upper left corner of the input feature map 310B, move the convolution kernel 313B one box at a time in the order from left to right and from top to bottom (or in the order from top to bottom and from left to right), and successively perform dot product operations on the weight values in the convolution kernel 313B and the feature values in the 3*3 region overlapping with the convolution kernel 313B in the input feature map 310B, then the feature values corresponding to all boxes in the output feature map 315B can be obtained.

[0128] Figure 3B It is used to illustrate the overlapping situation of the input data block in the up-down direction during convolution operation. As Figure 3BAs shown, when the convolution kernel is 3*3, in order to generate the output data block (hereinafter referred to as the upper output data block for convenience of description) composed of the regions with upper right diagonal lines in the output feature map 315B, it is necessary to use the input data block (hereinafter referred to as the upper input data block for convenience of description, with a size of 5*4, including the regions with upper right diagonal lines in the first and second rows of 310B and the regions with crossed lines in the third and fourth rows, that is, including the regions where the eigenvalue is located in the first five columns of each row in the first to fourth rows of 310B) composed of the regions with upper right diagonal lines and crossed lines (i.e., "", the same below) in the input feature map 310B to perform a convolution operation with the convolution kernel 313B. In order to generate the output data block (hereinafter referred to as the lower output data block for convenience of description) composed of the regions with lower right diagonal lines in the output feature map 315B, it is necessary to use the input data block (hereinafter referred to as the lower input data block for convenience of description, with a size of 5*4, including the regions with crossed lines in the third and fourth rows of 310B and the regions with lower right diagonal lines in the fifth and sixth rows, that is, including the regions where the eigenvalue is located in the first five columns of each row in the third to sixth rows of 310B) composed of the regions with lower right diagonal lines and crossed lines in the input feature map 310B to perform a convolution operation with the convolution kernel 313B. As Figure 3BAs shown, there is an overlapping area between the upper input data block and the lower input data block in the input feature map 310B. The overlapping area is the area with crosshatching in 310B. Specifically, when calculating the eigenvalue at (2, 1) in the output feature map (i.e., calculating the eigenvalue at the lower left corner of the upper output data block), the convolution kernel 313B needs to be used with the eigenvalues at (2, 1), (2, 2), (2, 3), (3, 1), (3, 2), (3, 3), (4, 1), (4, 2), and (4, 3) in the input feature map 310B; when calculating the eigenvalue at (3, 1) in the output feature map (i.e., calculating the eigenvalue at the upper left corner of the lower output data block), the convolution kernel 313B needs to be used with the eigenvalues at (3, 1), (3, 2), (3, 3), (4, 1), (4, 2), (4, 3), (5, 1), (5, 2), and (5, 3) in the input feature map 310B. Thus, when calculating the eigenvalue at the lower left corner of the upper output data block and the eigenvalue at the upper left corner of the lower output data block, the eigenvalues at (3, 1), (3, 2), (3, 3), (4, 1), (4, 2), and (4, 3) in the input feature map 310B will be used. Similarly, when calculating the eigenvalue at the lower right corner of the upper output data block and the eigenvalue at the upper right corner of the lower output data block, the eigenvalues at (3, 3), (3, 4), (3, 5), (4, 3), (4, 4), and (4, 5) in the input feature map 310B will be used. Since the eigenvalues at (3, 1), (3, 2), (3, 3), (3, 4), (3, 5), (4, 1), (4, 2), (4, 3), (4, 4), and (4, 5) in the input feature map 310B are used when calculating the eigenvalues in both the upper output data block and the lower output data block, this part of the area is called the overlapping area (i.e., the area with crosshatching in 310B). Therefore, when the convolution kernel is 3*3, there is an overlapping area of size 5*2 between the two input data blocks (i.e., the upper input data block and the lower input data block) in the input feature map 310B required to generate the two adjacent and non-overlapping output data blocks (i.e., the upper output data block and the lower output data block) in the output feature map 315B.

[0129] Now please refer to Figure 3C , Figure 3C which is a schematic diagram showing the overlapping situation of the input data block in the left-right direction when the convolution kernel is 3*3 during convolution operation according to an embodiment of the present invention. Figure 3C It is used to illustrate the overlapping situation of the input data block in the left-right direction during convolution operation. As Figure 3CAs shown, when the convolution kernel is 3*3, in order to generate the output data block with a right-upper diagonal line in the output feature map 315C (for ease of description, hereinafter referred to as the left output data block), it is necessary to use the input data block composed of the regions with right-upper diagonal lines and cross lines in the input feature map 310C (for ease of description, hereinafter referred to as the left input data block, with a size of 5*4, including the regions with right-upper diagonal lines in the 1st - 4th rows and the regions with cross lines in the 1st - 4th rows of 310C, that is, including the regions where the feature values are located in the first 5 columns of each row in the 1st - 4th rows of 310C) to perform a convolution operation with the convolution kernel 313C. In order to generate the output data block with a right-lower diagonal line in the output feature map 315C (for ease of description, hereinafter referred to as the right output data block), it is necessary to use the input data block composed of the regions with right-lower diagonal lines and cross lines in the input feature map 310C (for ease of description, hereinafter referred to as the right input data block, with a size of 5*4, including the regions with right-lower diagonal lines in the 1st - 4th rows and the regions with cross lines in the 1st - 4th rows of 310C, that is, including the regions where the feature values are located in the 4th - 8th columns of each row in the 1st - 4th rows of 310C) to perform a convolution operation with the convolution kernel 313C. As Figure 3C shown, there is an overlapping region between the left input data block and the right input data block in the input feature map 310C, and the overlapping region is the region with cross lines in 310C. Therefore, when the convolution kernel is 3*3, there is an overlapping region with a size of 2*4 between the two input data blocks (i.e., the left input data block and the right input data block) in the input feature map 310C that are needed to generate the two adjacent and non-overlapping output data blocks (i.e., the left output data block and the right output data block) in the output feature map 315C.

[0130] Now, please refer to Figure 3D , Figure 3D which is a schematic diagram showing the overlapping situation of the input data block in the upper-left to lower-right direction when performing a convolution operation with a 3*3 convolution kernel according to an embodiment of the present invention. Figure 3D It is used to illustrate the overlapping situation of the input data block in the upper-left to lower-right direction when performing a convolution operation. As Figure 3DAs shown, when the convolution kernel is 3*3, in order to generate the output data block with a right upper diagonal line in the output feature map 315D (for ease of description, hereinafter referred to as the upper left output data block), it is necessary to use the input data block composed of the regions with right upper diagonal lines and cross lines in the input feature map 310D (for ease of description, hereinafter referred to as the upper left input data block, with a size of 5*4, including the regions with right upper diagonal lines in the first to fourth rows and the regions with cross lines in the first to fourth rows of 310D, that is, including the regions where the feature values are located in the first 5 columns of each of the first to fourth rows of 310D) to perform a convolution operation with the convolution kernel 313D. In order to generate the output data block with a right lower diagonal line in the output feature map 315D (for ease of description, hereinafter referred to as the lower right output data block), it is necessary to use the input data block composed of the regions with right lower diagonal lines and cross lines in the input feature map 310D (for ease of description, hereinafter referred to as the lower right input data block, with a size of 5*4, including the regions with right lower diagonal lines in the first to fourth rows and the regions with cross lines in the first to fourth rows of 310D, that is, including the regions where the feature values are located in the fourth to eighth columns of each of the third to sixth rows of 310D) to perform a convolution operation with the convolution kernel 313D. As Figure 3D shown, there is an overlapping region between the upper left input data block and the lower right input data block in the input feature map 310D, and the overlapping region is the region with cross lines in 310D. Therefore, when the convolution kernel is 3*3, there is an overlapping region of size 2*2 between the two input data blocks (i.e., the upper left input data block and the lower right input data block) in the input feature map 310D that are needed to generate the two non-overlapping output data blocks (i.e., the upper left output data block and the lower right output data block) adjacent to each other in the upper left and lower right directions in the output feature map 315D.

[0131] Now please refer to Figure 3E , Figure 3E which is a schematic diagram showing the overlapping situation of the input data block in the lower left to upper right direction when the convolution kernel is 3*3 in accordance with an embodiment of the present invention. Figure 3E It is used to illustrate the overlapping situation of the input data block in the lower left to upper right direction during the convolution operation. As Figure 3EAs shown, when the convolution kernel is 3*3, in order to generate the output data block with a right-upper diagonal line in the output feature map 315E (for ease of description, hereinafter referred to as the lower-left output data block), it is necessary to use the input data block composed of the regions with right-upper diagonal lines and cross lines in the input feature map 310E (for ease of description, hereinafter referred to as the lower-left input data block, with a size of 5*4, including the regions with right-upper diagonal lines in the 1st - 4th rows and the regions with cross lines in the 1st - 4th rows of 310E, that is, including the regions where the feature values are located in the first 5 columns of each row in the 3rd - 6th rows of 310E) to perform a convolution operation with the convolution kernel 313E. In order to generate the output data block with a right-lower diagonal line in the output feature map 315E (for ease of description, hereinafter referred to as the upper-right output data block), it is necessary to use the input data block composed of the regions with right-lower diagonal lines and cross lines in the input feature map 310E (for ease of description, hereinafter referred to as the upper-right input data block, with a size of 5*4, including the regions with right-lower diagonal lines in the 1st - 4th rows and the regions with cross lines in the 1st - 4th rows of 310E, that is, including the regions where the feature values are located in the 4th - 8th columns of each row in the 1st - 4th rows of 310E) to perform a convolution operation with the convolution kernel 313E. As Figure 3E shown, there is an overlapping region between the lower-left input data block and the upper-right input data block in the input feature map 310E, and the overlapping region is the region with cross lines in 310E. Therefore, when the convolution kernel is 3*3, there is an overlapping region with a size of 2*2 between the two input data blocks (i.e., the lower-left input data block and the upper-right input data block) in the input feature map 310E that are needed to generate the two adjacent and non-overlapping output data blocks (i.e., the lower-left output data block and the upper-right output data block) in the output feature map 315E.

[0132] By Figure 3B-3EAnalysis shows that when the convolution kernel is 3×3, there is an overlapping area between the two input data blocks in the input feature map that are used to generate two adjacent and non-overlapping output data blocks in the generated output feature map. Similarly, when the convolution kernel is 5×5 or 7×7 (or a larger convolution kernel), there is also an overlapping area between the two input data blocks in the input feature map that are used to generate two adjacent and non-overlapping output data blocks in the generated output feature map. In addition, the larger the convolution kernel, the larger the overlapping area between the two input data blocks in the input feature map that are used to generate two adjacent and non-overlapping output data blocks in the generated output feature map. The width of the overlapping area between the two input data blocks in the input feature map that are used to generate two adjacent and non-overlapping output data blocks in the horizontal direction of the output feature map is the width of the convolution kernel minus the convolution stride in the horizontal direction (when the convolution kernel is 3×3 and the convolution stride in the horizontal direction is 1, the width of the overlapping area is 3 minus 1, that is, 2). The height of the overlapping area between the two input data blocks in the input feature map that are used to generate two adjacent and non-overlapping output data blocks in the vertical direction of the output feature map is the height of the convolution kernel minus the convolution stride in the vertical direction (when the convolution kernel is 3×3 and the convolution stride in the vertical direction is 1, the width of the overlapping area is 3 minus 1, that is, 2).

[0133] As can be seen from the above, during convolution operation, according to the width and height of the input data blocks that the convolution operation device can process in parallel, the input feature map is divided into multiple input data blocks. Assume that the size of the input data blocks that the convolution operation device can process in parallel is w×h (w is the width and h is the height, and both w and h are integers greater than 0), the convolution kernel is k×k (k is an integer greater than 0), and the convolution stride is s (s is an integer greater than 0). When k equals 1, there is no overlapping area between every two adjacent input data blocks (as shown in the situation in Figure 3A ); when k is greater than 1, there is an overlapping area between every two adjacent input data blocks, and the output data blocks generated after convolution operation on every two adjacent input data blocks are adjacent and non-overlapping (as shown in the situation in Figure 3B-3E ). Therefore, when the convolution kernel size and the convolution stride are known, the overlapping pattern between all input data blocks in the entire input feature map can be obtained. Figure 4 The input feature map after block division shown in Figure 3B-3E contains the overlapping situation between the input data blocks shown in Figure 4 . The following will elaborate on

[0134] Figure 4 is a schematic diagram of feature map block division when the convolution kernel is k×k (k is an integer greater than 0) and the convolution stride is s (s is an integer greater than 0) during convolution operation according to an embodiment of the present invention. As shown in Figure 4As shown, 410 is an input feature map of size W*H (both W and H are integers greater than 0), 413 is a convolution kernel of size k*k, and 415 is an output feature map generated after performing a blocked convolution operation on the input feature map 410. The size of the output feature map 415 is (W-(k-s))*(H-(k-s)), and the size of the output data block in the output feature map 415 is (w-(k-s))*(w-(k-s)). Figure 4 Among them, w is the width of the input data block (i.e., the width of the input data block that the convolution operation device can process in parallel), h is the height of the input data block (the height of the input data block that the convolution operation device can process in parallel), k is the side length of the convolution kernel, and s is the convolution step. The input feature map 410 is divided into multiple input data blocks with overlapping regions, such as input data blocks (1, 1), (1, 2), (1, 3), (2, 1), (2, 2), (2, 3), (3, 1), (3, 2), (3, 3)… etc. When k is greater than 1, there are overlapping regions between every two adjacent input data blocks, as shown Figure 3B to 3E As shown. When there are overlapping regions in the input data block, these overlapping regions can be further classified. For example, the input data block (1, 1) in the input feature map 410 contains 4 regions: non-overlapping region E 1,1 、right vertical overlapping region F 1,1 、bottom horizontal overlapping region H 1,1 and bottom right overlapping region T 1,1 ; among them, the right vertical overlapping region F 1,1 of the input data block (1, 1) is also the left vertical overlapping region of the input data block (1, 2), and the bottom horizontal overlapping region H 1,1 of the input data block (1, 1) is also the top horizontal overlapping region of the input data block (2, 1). The bottom right overlapping region T 1,1 of the input data block (1, 1) is also the bottom left overlapping region of the input data block (1, 2), the top right overlapping region of the input data block (2, 1), and the top left overlapping region of the input data block (2, 2). The input data block (2, 2) contains 9 regions: non-overlapping region E 2,2 、right vertical overlapping region F 2,2 、bottom horizontal overlapping region H 2,2 、bottom right overlapping region T 2,2 、top left overlapping region T 1,1 、top horizontal overlapping region H 1,2 、top right overlapping region T 1,2 、left vertical overlapping region F 2,1 and bottom left overlapping region T 2,1 ; among them, the top left overlapping region T 1,1Also the lower right overlapping area of the input data block (1, 1) and the upper horizontal overlapping area H of the input data block (2, 2) 1,2 Also the lower horizontal overlapping area of the input data block (1, 2) and the upper right overlapping area T of the input data block (2, 2) 1,2 Also the lower left overlapping area of the input data block (1, 3) and the left vertical overlapping area F of the input data block (2, 2) 2,1 Also the right vertical overlapping area of the input data block (2, 1) and the right vertical overlapping area F of the input data block (2, 2) 2,2 Also the left vertical overlapping area of the input data block (2, 3) and the lower left overlapping area T of the input data block (2, 2) 2,1 Also the upper right overlapping area of the input data block (3, 1) and the lower horizontal overlapping area H of the input data block (2, 2) 2,2 Also the upper horizontal overlapping area of the input data block (3, 2) and the lower right overlapping area T of the input data block (2, 2) 2,2 Also the upper left overlapping area of the input data block (3, 3). Obviously, the overlapping methods of all input data blocks can be represented by the non-overlapping area E x,y 、the left (right) vertical overlapping area F x,y 、the upper (lower) horizontal overlapping area H x,y 、and the lower left (upper left / upper right / lower right) corner overlapping area T x,y and will not be elaborated here.

[0135] Such as Figure 4As shown, each input data block in the input feature map 410 contains at most 9 regions, while the input data blocks located in the first row, first column, last row, and last column of the input feature map 410 contain fewer than 9 regions. Specifically, the input data block in the first row of the input feature map 410 does not contain the upper-left overlapping region, upper-horizontal overlapping region, and upper-right overlapping region; the input data block in the first column of the input feature map 410 does not contain the upper-left overlapping region, left-vertical overlapping region, and lower-left overlapping region; the input data block in the last row of the input feature map 410 does not contain the lower-left overlapping region, lower-horizontal overlapping region, and lower-right overlapping region; the input data block in the last column of the input feature map 410 does not contain the upper-right overlapping region, right-vertical overlapping region, and lower-right overlapping region. For example, the input data block (1, 1) contains 4 regions, and the input data block (3, 1) contains 6 regions. For the convenience of subsequent description, we regard all input data blocks in the input feature map as input data blocks containing 9 regions. For a specific input data block, if it does not contain certain regions, we regard it as an input data block containing these regions, and regard the size of these regions as 0*0 (that is, both the width and height are 0). For example, we regard the input data block (3, 1) in the input feature map as an input data block with left-vertical overlapping region, upper-left overlapping region, and lower-right overlapping region having a size of 0*0.

[0136] In another embodiment, the convolution kernel is rectangular, with its length represented by k1 and its width represented by k2 (k1 and k2 can be integers greater than 0 and k1 is not equal to k2). Different from Figure 4 the embodiment where the convolution kernel shown is square is that: the width of the horizontal overlapping region between the input data blocks (1, 1) and (1, 2) is k1 - s, and the height of the vertical overlapping region between the input data blocks (1, 1) and (2, 1) is k2 - s. The size of the output feature map 415 is (W - (k1 - s)) * (H - (k2 - s)), and the size of the output data block in the output feature map 415 is (w - (k1 - s)) * (h - (k2 - s)); other aspects are the same as the embodiment where the convolution kernel is square.

[0137] In another embodiment, when performing the convolution operation, different convolution strides can be used horizontally and vertically. For example, the horizontal convolution stride is s1 and the vertical convolution stride is s2 (s1 and s2 can be integers greater than 0). Different from Figure 4The difference between the embodiments with a horizontal convolution stride and a vertical convolution stride both being s is as follows: the width of the horizontal overlapping region between the input data blocks (1, 1) and (1, 2) is k - s1, the height of the vertical overlapping region between the input data blocks (1, 1) and (2, 1) is k - s2, and the size of the output data block in the output feature map 415 is (w - (k - s1)) * (h - (k - s2)); in other aspects, it is the same as the embodiment with a horizontal convolution stride and a vertical convolution stride both being s. In another embodiment, the convolution kernel is a rectangle with a length of k1 and a width of k2, and the horizontal and vertical convolution strides are different s1 and s2 (k1, k2, s1, and s2 are all integers greater than 0), so the size of the output data block in the output feature map 415 is (w - (k1 - s1)) * (h - (k2 - s2)).

[0138] In the following description of the present disclosure, for the input feature map that needs to be divided into blocks for convolution operations (in the case where the width and height of the input feature map are both smaller than the width and height of the input data block that the convolution operation module can process in parallel, the convolution operation module can directly process one input feature map at a time, so there is no need to divide the input feature map into blocks), it is all divided into multiple input data blocks with overlapping regions in the manner Figure 4 shown, and then all the input data blocks are convolved with the convolution kernel in the order from left to right and from top to bottom (or from top to bottom and from left to right) to generate the corresponding output data blocks in the output feature map. The generated output data blocks are combined in the order from left to right and from top to bottom (or from top to bottom and from left to right) to generate the output feature map.

[0139] In addition, for the convenience of describing the processing flow of the input feature map in the order from left to right and from top to bottom in the following text, we divide each input data block in the input feature map 410 into three parts: the horizontal main region, the upper horizontal sub-region, and the lower horizontal sub-region. Specifically, we combine the non-overlapping region, the left vertical overlapping region, and the right vertical overlapping region of each input data block in the input feature map 410 into the horizontal main region. For example, the horizontal main region of the input data block (1, 1) is: E 1,1 +F 1,1 ; the horizontal main region of the input data block (1, 2) is: F 1,1 +E 1,2 +F 1,2 ; the horizontal main region of the input data block (2, 2) is: F 2,1 +E 2,2 +F 2,2 . We combine the lower left overlapping region, the lower horizontal overlapping region, and the lower right overlapping region of each input data block in the input feature map 410 into the lower horizontal sub-region. For example, the lower horizontal sub-region of the input data block (1, 1) is: H 1,1+T 1,1 For the lower horizontal sub-region of the input data block (1, 2), it is: T 1,1 +H 1,2 +T 1,2 For the lower horizontal sub-region of the input data block (2, 2), it is: T 2,1 +H 2,2 +T 2,2 We collectively refer to the upper left overlapping region, upper horizontal overlapping region, and upper right overlapping region of each input data block in the input feature map 410 as the upper horizontal sub-region. For example, the upper horizontal sub-region of the input data block (3, 1) is: H 2,1 +T 2,1 For the upper horizontal sub-region of the input data block (3, 2), it is: T 2,1 +H 2,2 +T 2,2 For the upper horizontal sub-region of the input data block (3, 3), it is: T 2,2 +H 2,3 +T 2,3 The sizes of the upper horizontal sub-regions of the input data blocks (1, 1), (1, 2), and (1, 3) are all 0*0. We collectively refer to all the lower horizontal overlapping regions and lower right overlapping regions of each row of input data blocks in the input feature map 410 as the lower horizontal row overlapping region. For example, the lower horizontal row overlapping region of the input data blocks in the first row is: H 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 +…. We collectively refer to all the upper horizontal overlapping regions and upper right overlapping regions of each row of input data blocks in the input feature map 410 as the upper horizontal row overlapping region. For example, the upper horizontal row overlapping region of the input data blocks in the third row (which is also the lower horizontal row overlapping region of the input data blocks in the second row) is: H 2,1 +T 2,1 +H 2,2 +T 2,2 +H 2,3 +T 2,3 +…. The size of the upper horizontal row overlapping region of the input data blocks in the first row is 0*0.

[0140] Similarly, for the convenience of describing the processing flow of the input feature map in the order from top to bottom and from left to right (i.e., processing column by column) in the following text, we collectively refer to the non-overlapping region, lower horizontal overlapping region, and upper horizontal overlapping region of each input data block in the input feature map 410 as the vertical main region. For example, the vertical main region of the input data block (1, 1) is: E 1,1 +H 1,1 For the vertical main region of the input data block (2, 1), it is: H 1,1+E 2,1 +H 2,1 For the vertical main area of the input data block (2, 2): H 1,2 +E 2,2 +H 2,2 We collectively refer to the upper-left overlapping area, left vertical overlapping area, and lower-left overlapping area of each input data block in the input feature map 410 as the left vertical secondary area. For example, the left vertical secondary area of the input data block (1, 3) is: F 1,2 +T 1,2 For the left vertical secondary area of the input data block (2, 3): T 1,2 +F 2,2 +T 2,2 For the left vertical secondary area of the input data block (3, 3): T 2,2 +F 3,2 +T 3,2 We collectively refer to the upper-right overlapping area, right vertical overlapping area, and lower-right overlapping area of each input data block in the input feature map 410 as the right vertical secondary area. For example, the right vertical secondary area of the input data block (1, 3) is: F 1,3 +T 1,3 For the right vertical secondary area of the input data block (2, 3): T 1,3 +F 2,3 +T 2,3 For the right vertical secondary area of the input data block (3, 3): T 2,3 +F 3,3 +T 3,3 The sizes of the left vertical secondary areas of the input data blocks (1, 1), (2, 1), and (3, 1) are all 0*0. We collectively refer to the right vertical overlapping area and lower-right overlapping area of each column of input data blocks in the input feature map 410 as the right vertical column overlapping area. For example, the right vertical column overlapping area of the first column is: F 1,1 +T 1,1 +F 2,1 +T 2,1 +F 3,1 +T 3,1 +… We collectively refer to the left vertical overlapping area and lower-left overlapping area of each column of input data blocks in the input feature map 410 as the left vertical column overlapping area. For example, the left vertical column overlapping area of the third column (which is also the right vertical column overlapping area of the second column) is: F 1,2 +T 1,2 +F 2,2 +T 2,2 +F 3,2 +T 3,2+…。For ease of description, hereinafter, the horizontal main region and the vertical main region are referred to as the main regions, the lower horizontal sub-region and the right vertical sub-region are referred to as the first sub-regions, the lower left overlapping region and the upper right overlapping region of the input data block are referred to as the first overlapping sub-regions of the first sub-regions, the lower horizontal overlapping region and the right vertical overlapping region of the input data block are referred to as the second overlapping sub-regions of the first sub-regions, the lower right overlapping region of the input data block is referred to as the third overlapping sub-region of the first sub-regions; the first overlapping sub-regions, the second overlapping sub-regions and the third overlapping sub-regions are referred to as overlapping sub-regions; the upper horizontal sub-region and the left vertical sub-region are referred to as the second sub-regions; the first sub-regions and the second sub-regions are referred to as sub-regions.

[0141] Through Figure 4 the input feature map 410 and its related description in, it can be known that each sub-region of the input data block contains at least one overlapping sub-region, wherein the number of input data blocks adjacent to the overlapping sub-region of the sub-region of the input data block is greater than the number of input data blocks adjacent to the overlapping region of the main region of the input data block.

[0142] Figure 5 FIG. 9 is a block diagram of a computing device 500 including a convolution operation module 530 according to an embodiment of the present invention. In one embodiment, the computing device 500 is, for example, a server, a desktop computer, a notebook computer, a mobile phone, a tablet, or other electronic devices with computing functions.

[0143] As Figure 5 shown, the computing device 500 includes a memory 520 and a convolution operation module 530. The memory 520 is coupled to the convolution operation module 530. The convolution operation module 530 can be used to perform the convolution operation of the convolution layer in a convolutional neural network (for example, Figure 1 the convolutional neural network 100 shown). The memory 520 is used to store the input feature map set of the current convolution layer in the convolutional neural network, the output feature map set of the current convolution layer, the parameters of each convolution layer, and the convolution kernel group set of each convolution layer. The current convolution layer refers to the convolution layer that the convolution operation module 530 is processing or about to process. In one embodiment, the memory 520 is a system memory. In another embodiment, the memory 520 is a static random access memory (SRAM). In another embodiment, the memory 520 can be any memory used by the computing device 500 to store data.

[0144] As Figure 5As shown, the convolution operation module 530 includes a configuration register 531, a secondary processing module 538, a cache 532, a primary processing module 534, an arithmetic unit 536, and a data processing module 539. The secondary processing module 538 is coupled to the cache 532 and is used to read the input feature map and the convolution kernel from the memory 520, then perform secondary decompression on the read input feature map to generate a primary compressed input feature map, and then store the primary compressed input feature map and the convolution kernel in the cache 532. The primary processing module 534 is coupled to the cache 532 and the arithmetic unit 536 and is used to read the primary compressed input feature map and the convolution kernel from the cache 532, perform primary decompression on the primary compressed input feature map to generate the original data of the input feature map (i.e., uncompressed data), and then send the input feature map and the convolution kernel to the arithmetic unit 536 for convolution operation. The arithmetic unit 536 is coupled to the primary data processing module 534 and the data processing module 539 and is used to receive the input feature map and the convolution kernel sent by the primary processing module 534, perform convolution operation on the received input feature map and the convolution kernel to generate an output feature map, and then send the output feature map to the data processing module. The data processing module 539 includes a segmentation module 535 and a compression module 537. The segmentation module 535 receives the output feature map generated by the arithmetic unit 536 and segments the output feature map into multiple output data blocks; then, the compression module 537 performs two-level compression on the multiple output data blocks and stores them in the memory 520. The configuration register 531 is used to store the parameters of the current convolutional layer (the usage of these parameters will be described later). The cache 532 includes a cache segment 5321 and a cache segment 5323. The cache segment 5321 is used to cache the input feature map data of the current convolutional layer, and the cache segment 5323 is used to cache the convolution kernel group of the current convolutional layer. The arithmetic unit 536 includes multiple arithmetic units (arithmetic units 5361 to 536Z), and each arithmetic unit can perform convolution operation on an input data block and a convolution kernel to generate an output data block. In the present disclosure, it is assumed that the size of the input data block that each arithmetic unit of the arithmetic unit 536 can process is w*h. The following describes the processing flow of the convolution operation module 530 for performing the convolution operation of the current convolutional layer. The parameters written into the configuration register 531 include the address of the input feature map set of the current convolutional layer (i.e., the first convolutional layer) in the memory 520, the address of the output feature map set of the current convolutional layer in the memory 520, the width and height of the input feature map of the current convolutional layer, the address of the convolution kernel group set of the current convolutional layer in the memory 520, the width and height of the convolution kernel in the convolution kernel group of the current convolutional layer, the convolution stride of the current convolutional layer, the padding size (padding) of the current convolutional layer, the width and height of the convolution kernel in the convolution kernel group of the next convolutional layer, and the padding size of the next convolutional layer.Among them, the width and height of the input feature map of the current convolutional layer of the parameter, the address of the set of convolutional kernel groups of the current convolutional layer in the memory 520, the width and height of the convolutional kernels in the convolutional kernel group of the current convolutional layer, the convolutional stride of the current convolutional layer, the padding size (padding) of the current convolutional layer, the width and height of the convolutional kernels in the convolutional kernel group of the next convolutional layer, and the padding size of the next convolutional layer are read from the storage segment 525 of the memory 520.

[0145] First, the secondary processing module 538 reads the input feature map of the current convolutional layer from the memory 520 according to the parameters in the configuration register 531 (the input feature map stored in the memory 520 has been compressed twice, and the processing flow of storing the input feature map in the memory 520 through two - level compression will be described in detail later), and performs secondary decompression on it to obtain the first - level compressed data of the input feature map of the current convolutional layer, and then stores the first - level compressed data of the input feature map of the current convolutional layer in the cache segment 5321 of the cache 532. On the other hand, the secondary processing module 538 also reads the convolutional kernel group of the current convolutional layer from the memory 520 according to the parameters in the configuration register 531, and stores it in the cache segment 5323 of the cache 532.

[0146] Then, the primary processing module 534 reads the first - level compressed data of the input feature map of the current convolutional layer from the cache segment 5321, and performs primary decompression on it (see the format of the first - level compressed data above) to obtain the input feature map of the current convolutional layer. The primary processing module 534 also reads the convolutional kernel group corresponding to the input feature map of the current convolutional layer from the cache segment 5323. Then the primary processing module 534 sends the input feature map of the current convolutional layer and the convolutional kernels in the corresponding convolutional kernel group to the arithmetic unit 536 for convolution operation.

[0147] Then, the arithmetic unit 536 will allocate the input feature map of the current convolutional layer and the corresponding convolutional kernels to the idle arithmetic units for convolution operation according to the parameters in the configuration register 531 to generate the output feature map, and send the generated output feature map to the data processing module 539.

[0148] Finally, the data processing module 539 performs two - level compression on the received output feature map according to the parameters in the configuration register 531 (the processing flow of the two - level compression will be described in detail later), and then writes it into the memory 520. The output feature map of the current convolutional layer will be used as the input feature map of the next convolutional layer and participate in the convolutional operation of the next convolutional layer. Since the input feature map of the first convolutional layer is the original input data of the convolutional operation, it needs to be compressed at two levels and then stored in the memory 520 before the convolutional operation is performed by the computing device 500. In one embodiment, the convolutional operation module 530 also provides a decompression / compression interface externally. Through this decompression / compression interface, a module outside the convolutional operation module 530 can call the data processing module 539 to perform a compression operation, or call the secondary processing module 538 and / or the primary processing module 534 to perform a decompression operation; at this time, the data processing module 539, the secondary processing module 538, and the primary processing module 534 are only simply called. The computing device 500 can perform two - level compression on the input feature map of the first convolutional layer through the decompression / compression interface provided by the convolutional operation module 530 and then store it in the memory 520.

[0149] In another embodiment, the secondary processing module 538, the cache 532, the primary processing module 534, the arithmetic unit 536, and the data processing module 539 can be implemented in a pipeline to improve the processing speed of the convolutional operation module 530.

[0150] As described above, during the convolutional operation, many elements with a value of 0 will be generated in the input feature map / output feature map. Therefore, through the primary compression of the present invention, a large amount of data required for the convolutional operation can be compressed, so the space required when stored in the cache 532 will be greatly reduced. In addition, since there are many levels of convolutional operations, through the two - level compression of the present invention, the input feature map / output feature map of each convolutional layer can be effectively compressed, which can greatly reduce the data transmission volume between the convolutional operation module 530 and the memory 520 (because of the two - level compression), thereby improving the overall operation efficiency of the computing device 500. In addition, when the input feature map is sent to the convolutional operation module 530 for processing, since the arithmetic unit 536 cannot process compressed data (it can only process the original data of the input feature map), the primary compressed data of the input feature map is stored in the cache 532, and it is decompressed by the primary decompression module 534 before the input feature map is sent to the arithmetic unit 536 for processing.

[0151] Figure 6A FIG. 10 is a schematic diagram of the data stored in the memory 520 of the computing device 500 according to an embodiment of the present invention. Figure 6B FIG. 12 is a more detailed block diagram of the computing device 500 according to an embodiment of the present invention. Figure 6CThe processing flow of compressing the input feature map of the Nth convolutional layer in two levels and then writing it into the memory according to an embodiment of the present invention Figure 6D The processing flow of the computing device 500 generating an output feature map according to an embodiment of the present invention Figure 6E The processing flow of the computing device 500 generating an output feature map according to another embodiment of the present invention Figure 6F-1 to 6F-2 A more detailed processing flow of the computing device 500 generating an output feature map according to an embodiment of the present invention. The following will be combined with Figure 6A 、 6B 、6C, 6D, 6E and 6F-1 to 6F-2 to introduce in detail the processing flow of running a convolutional neural network using the convolutional operation device 500

[0152] As Figure 6A shown, the memory 520 includes storage segments 521, 523, 525 and 527. The memory 520 is used to store the data required for running the convolutional neural network. For example, the storage segment 521 is used to store the set of input feature maps of the current convolutional layer, the storage segment 523 is used to store the set of output feature maps of the current convolutional layer (before performing the convolutional operation of the current convolutional layer, the number of output feature maps stored in the storage segment 523 is 0), the storage segment 525 is used to store the parameters of all convolutional layers, and the storage segment 527 is used to store the set of convolutional kernel groups of all convolutional layers. The storage segment 525 is used to store the parameters related to each convolutional layer. For example, the parameters related to the first convolutional layer include: the width and height of the input feature map of the first convolutional layer, the address of the set of convolutional kernel groups of the first convolutional layer in the memory 520, the width and height of the convolutional kernels in the convolutional kernel group of the first convolutional layer, the convolutional stride of the first convolutional layer, and the padding size of the first convolutional layer. The parameters of other convolutional layers in the storage segment 525 are similar to those of the first convolutional layer, and will not be elaborated here. It should be noted that before the convolutional operation starts, the parameters and the set of convolutional kernel groups related to each convolutional layer will be separately stored in the storage segment 525 and the storage segment 527, and will not change during the convolutional operation

[0153] Before using the computing device 500 to run the convolutional neural network, it is necessary to first store the data required for running the convolutional neural network in the memory 520. Specifically, the computing device 500 writes the parameters of the first to X convolutional layers into the storage segment 525, writes the set of convolutional kernel groups of the first to X convolutional layers into the storage segment 527, and writes the set of input feature maps of the first convolutional layer according to Figure 6CThe processing flow in [the device] is written into the storage segment 521 after two - stage compression. At this time, since the first convolution operation has not started yet, the output feature map of the first convolution layer has not been generated, so there is no output feature map stored in the storage segment 523. It should be noted that only the input feature map set of the first convolution layer is written into the memory 520 by the computing device 500 through the compression interface called by the convolution operation module 530; the input feature map sets of other convolution layers are the output feature map sets of the previous convolution layer, which are directly compressed in two - stages and stored in the memory 520 after being received by the data processing module 539. For example, the output feature map set of the first convolution layer is the input feature map set of the second convolution layer, and the output feature map set of the first convolution layer is written into the storage segment 523 by the data processing module 539 (after two - stage compression). The data processing module 539 writes the output feature map set of the current convolution layer into the storage segment 523 through Figure 6C the processing flow. Next, the processing flow of writing all the input feature maps of the N - th convolution layer into the memory after two - stage compression will be described in detail with reference to Figure 6C .

[0154] As Figure 6C shown, in step S601C, the segmentation module 535 generates input data blocks. Specifically, the segmentation module 535 in the data processing module 539 divides all the input feature maps of the N - th convolution layer into input data blocks with overlapping regions according to the width and height of the input data blocks that the convolution operation device 530 can process in parallel, the width and height of the convolution kernel of the N - th convolution layer, and the convolution stride of the N - th convolution layer (these parameters can be obtained from the configuration register 531) (using the segmentation method as shown in Figure 4 ). Then step S603C is executed.

[0155] In step S603C, the compression module 537 performs first - stage compression on the input data blocks. Specifically, the compression module 537 in the data processing module 539 compresses the main region of each input data block of the input feature map (for example, when processing the input data blocks in the order from left to right and from top to bottom, the main region of the input data block (2, 2) is F 2,1 +E 2,2 +F 2,2 ; when processing the input data blocks in the order from top to bottom and from left to right, the main region of the input data block (2, 2) is H 1,2 +E 2,2 +H 2,2 ) and the secondary region (for example, when processing the input data blocks in the order from left to right and from top to bottom, the first secondary region of the input data block (2, 2) is: T 2,1 +H 2,2 +T 2,2; When processing the input data blocks in the order from top to bottom and from left to right, the first region of the input data block (2, 2) is T 1,2 +F 2,2 +T 2,2 ), respectively perform primary compression to generate the primary region and secondary region after primary compression. In another embodiment, when processing the input data blocks in the order from left to right and from top to bottom, the first regions of all the input data blocks on the same row (for example, the first regions of all the input data blocks on the second row are H 2,1 +T 2,1 +H 2,2 +T 2,2 +H 2,3 +T 2,3 +… It should be noted that the first regions H of all the input data blocks on the first row 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 +… are also regarded as a whole for primary compression at the same time as the second regions of all the input data blocks on the second row); Similarly, when processing the input data blocks in the order from top to bottom and from left to right, the first regions of all the input data blocks on the same column (for example, the first regions of all the input data blocks on the second column are F 1,2 +T 1,2 +F 2,2 +T 2,2 +F 3,2 +T 3,2 +… It should be noted that the first regions F of all the input data blocks on the first column 1,1 +T 1,1 +F 2,1 +T 2,1 +F 3,1 +T 3,1 +… are also regarded as a whole for primary compression at the same time as the second regions of all the input data blocks on the second column). Then step S605C is executed.

[0156] In step S605C, the compression module 537 performs secondary compression on the input data blocks after primary compression. Specifically, the compression module 537 in the data processing module 539 performs secondary compression on the primary region and secondary region after primary compression of each input data block of the input feature map, respectively, to generate the primary region and secondary region after two - level compression. In another embodiment, the primary regions of multiple (such as 5) adjacent input data blocks in the same input feature map can be regarded as a whole (for example, connected in sequence one after another) for secondary compression. Then step S607C is executed.

[0157] In step S607C, the data processing module 539 stores the input data block after secondary compression into the memory 520. Specifically, the data processing module 539 stores the primary region and the secondary region after secondary compression of each input data block of the input feature map into the storage segment 521 (for example, stores the input feature map of the first convolutional layer into the storage segment 521) or the storage segment 523 (for example, stores the input feature map of the second convolutional layer into the storage segment 523, that is, stores the output feature map of the first convolutional layer into the storage segment 523) of the memory 520.

[0158] Now let's go back to Figure 6A . As Figure 6A shown, before performing the convolution operation of the current convolutional layer, all input feature maps (input feature maps 5211 to input feature map 521M) of the current convolutional layer are sequentially stored in the storage segment 521, and for each input feature map, its primary region is stored first, and then its secondary region is stored. For example, when storing the input feature map 5211, first store all the primary regions of the input feature map 5211 into the primary region 52111 of the input feature Figure 1 in the storage segment 521 in the order from left to right and from top to bottom, and then store all the secondary regions of the input feature Figure 1 in the lower horizontal row overlapping region 52112 of the input feature Figure 1 in the order from left to right and from top to bottom. Taking the input feature map 410 in Figure 4 (assuming the input feature map 410 is the input feature Figure 1 ) as an example, when storing the input feature map 410, first sequentially store the primary regions E 1,1 +F 1,1 of the input data block (1, 1) of the input feature map 410, the primary regions F 1,1 +E 1,2 +F 1,2 , etc. of the input data block (1, 2) into the primary region 52111 of the input feature Figure 1 in the storage segment 521. Then, sequentially store the first secondary regions of the input data blocks in the first row of the input feature map 410, the first secondary regions of the input data blocks in the second row... etc. in the secondary region 52112 of the input feature Figure 1 . The way of storing the output feature map in the storage segment 523 is the same as the way of storing the input feature map in the storage segment 521, which will not be elaborated here.

[0159] In another embodiment, when storing the input feature map (or output feature map) in the storage segment 521 (or storage segment 523), the secondary region is stored first, and then the primary region is stored after the secondary region.

[0160] After writing the input feature map set of the first convolutional layer into the memory 520 through two-level compression, the arithmetic unit 500 first writes the parameters of the first convolutional layer into the configuration register 531, and then notifies the convolutional arithmetic module 530 to start the convolutional operation of the first convolutional layer.

[0161] After receiving the notification to start the convolutional operation, the computing device 500 will Figure 6D or Figure 6E (which will be described in detail later) perform a convolutional operation on the input feature map set of the first convolutional layer and each convolutional kernel group to generate an output feature map corresponding to each convolutional kernel group. First, the processing flow of Figure 6D for performing a convolutional operation on the input feature map set and a convolutional kernel group to generate an output feature map will be described. The computing device 500 first executes step S603D.

[0162] In step S603D, each of the multiple input data blocks is divided into multiple non-overlapping regions, where there is an overlapping region between any two adjacent input data blocks. Specifically, the input feature map is divided into multiple input data blocks, where there is an overlapping region between any two adjacent input data blocks; according to the overlapping regions between the input data blocks, each input data block is divided into multiple non-overlapping regions. Specifically, the computing device 500 uses the processing flow in Figure 6C step S601C described above to divide the input feature map into multiple input data blocks with overlapping regions. Then, the computing device 500 divides each input data block into multiple non-overlapping regions according to the overlapping regions between the input data blocks, that is, each input data block is divided into a main region, a first secondary region, and a second secondary region. As Figure 4 shown, when processing the input feature map in the order from left to right and top to bottom, the input data block (2, 2) is divided into a main region (F 2,1 +E 2,2 +F 2,2 ), a first secondary region (T 2,1 +H 2,2 +T 2,2 ), and a second secondary region (T 1,1 +H 1,2 +T 1,2 ), the input data block (1, 2) is divided into a main region (F 1,1 +E 1,2 +F 1,2 ), a first secondary region (i.e., the second secondary region of the input data block (2, 2), T 1,1 +H 1,2 +T 1,2) The input data block (2, 2) is adjacent to (1, 2), and there is an overlapping area T between the input data block (2, 2) and (1, 2). 1,1 +H 1,2 +T 1,2 ; When processing the input feature map in the order from top to bottom and from left to right, the input data block (2, 2) is divided into a main area (H 1,2 +E 2,2 +H 2,2 ), a first area (T 1,2 +F 2,2 +T 2,2 ), and a second area (T 1,1 +F 2,1 +T 2,1 ). The input data block (2, 1) is divided into a main area (H 1,1 +E 2,1 +H 2,1 ), a first area (i.e., the second area of the input data block (2, 2), T 1,1 +F 2,1 +T 2,1 ). The input data block (2, 1) is adjacent to (2, 2), and there is an overlapping area T between the input data block (2, 1) and (2, 2). 1,1 +F 2,1 +T 2,1 . Then, the computing device 500 stores the regions of each input data block of the input feature map in the memory 520 after two - level compression according to Figure 6C the steps S603C, S605C, and S607C in. Then, step S605D is executed.

[0163] In step S605D, the computing device 500 stores the multiple non - overlapping regions of each of the input data blocks in the corresponding non - overlapping storage spaces in the cache. Specifically, the computing device 500 reads the regions of the input data blocks that have been compressed at two levels from the memory 520, decompresses them at the second level, and stores them in the cache 532. For a more detailed process, refer to the description of Figure 6F-1 to 6F-2 the steps S603F, S605F, S607F, and S609F in. Then, step S607D is executed.

[0164] In step S607D, the computing device 500 generates each of the input data blocks according to the regions corresponding to each of the input data blocks stored in the non - overlapping storage spaces. Specifically, the computing device 500 generates the corresponding input data blocks according to the regions of the input data blocks that have been compressed at the first level and stored in the cache 532. For a more detailed process, refer to the description of Figure 6F-1 to 6F-2Descriptions of steps S613F, S615F, S617F, and S619F. Then step S609D is executed.

[0165] In step S609D, the computing device 500 performs a convolution operation on the generated multiple input data blocks to generate the output feature map. Specifically, the computing device 500 sends the input data blocks to the arithmetic unit 536 for convolution operation to generate output data blocks, and then splices the output data blocks into an output feature map. For a more detailed process, refer to the descriptions of Figure 6F-1 to 6F-2 steps S621F, S623F, S625F, S627F, and S629F.

[0166] Through the above descriptions of Figure 6C and Figure 6D it can be known that the input data blocks stored in the memory 520 are data that have been compressed at the first level and then compressed at the second level, and the input data blocks stored in the cache 532 are data that have been compressed at the first level. Among them, the compression ratio of the input data blocks stored in the memory 520 is higher than that of the input data blocks stored in the cache 532. Therefore, when the convolution operation module 530 loads data from the external memory 520 or transfers data to the memory 520 for storage by the convolution operation module 530, both the required data transfer volume and transfer time can be significantly reduced, thus improving the execution efficiency of the system.

[0167] Next, the processing flow of convolving an input feature map set with a set of convolution kernels to generate an output feature map in Figure 6E is described. The computing device 500 first executes step S601E.

[0168] In step S601E, the computing device 500 performs a second-level decompression operation on the input feature map, where the input feature map includes multiple input data blocks and there is an overlapping area between any two adjacent input data blocks, and each input data block includes a main area and at least one sub-area. Specifically, the computing device 500 reads the input data block area of the input feature map from the memory 520, and then performs a second-level decompression operation on the read input data block area. For a more detailed process, refer to the descriptions of Figure 6F-1 to 6F-2 steps S603F, S605F, S607F, and S609F. Then step S603E is executed.

[0169] In step S603E, the computing device 500 stores the primary region after the secondary decompression operation and at least one secondary region after the secondary decompression operation of each of the input data blocks into different storage spaces respectively. Specifically, the computing device 500 stores the primary region after the secondary decompression operation and at least one secondary region after the secondary decompression operation of each of the input data blocks into different storage spaces in the cache 532. For a more detailed process, refer to the descriptions of steps S603F, S605F, S607F, and S609F of Figure 6F-1 to 6F-2 below. Then, step S605E is executed.

[0170] In step S605E, the computing device 500 performs a primary decompression operation on the primary region after the secondary decompression operation and at least one secondary region after the secondary decompression operation of each of the input data blocks. Specifically, the computing device 500 reads the primarily compressed primary region and secondary region of the input data block from the cache 532, performs a primary decompression operation on the read primarily compressed primary region and secondary region, and stores the result in the register 5342. For a more detailed process, refer to the description of step S613F of Figure 6F-1 to 6F-2 below. Then, step S607E is executed.

[0171] In step S607E, the computing device 500 generates each of the input data blocks by using the primary region after the primary decompression operation and the secondary region after the primary decompression operation of each of the input data blocks. Specifically, the computing device 500 reads the primary region and secondary region after the primary decompression operation of the input data block from the register 5342 to generate the input data block. For a more detailed process, refer to the description of step S619F of Figure 6F-1 to 6F-2 below. Then, step S609E is executed.

[0172] In step S609E, the computing device 500 performs a convolution operation on each of the input data blocks to generate the output feature map. Specifically, the computing device 500 sends the input data block to the arithmetic unit 536 for convolution operation to generate an output data block, and then splices the output data blocks into an output feature map. For a more detailed process, refer to the descriptions of steps S621F, S623F, S625F, S627F, and S629F of Figure 6F-1 to 6F-2 below.

[0173] The following describes Figure 6F-1 to 6F-2 a more detailed processing flow of convolving an input feature map set with a convolution kernel group to generate an output feature map. The convolution operation module 530 first executes step S601F.

[0174] In step S601F, the secondary processing module 538 reads a set of convolution kernels of the current convolutional layer from the memory and stores them in the cache 532. Specifically, the secondary processing module 538 reads an unprocessed set of convolution kernels of the current convolutional layer from the storage segment 527 of the memory 520 according to the address of the set of convolution kernels of the current convolutional layer stored in the configuration register 531, and stores them in the cache segment 5323 of the cache 532. According to the description of the present disclosure for Figure 2 , each set of convolution kernels may include multiple convolution kernels (such as convolution kernels 1 to convolution kernel M shown in the cache segment 5323). Then step S603F is executed.

[0175] In step S603F, the secondary processing module 538 reads the two-stage compressed main regions of the input data blocks located at the same position in all the input feature maps from the memory 520 (for example, the two-stage compressed main region of the input data block (1, 1) in all the input feature maps; when processing the input data blocks in the order from left to right and top to bottom, the main region refers to the horizontal main region; when processing the input data blocks in the order from top to bottom and left to right, the main region refers to the vertical main region; the same below). Specifically, the secondary processing module 538 reads a two-stage compressed main region located at the same position of each input feature map from the storage segment 521 of the memory 520 according to the address of the set of all the input feature maps of the current convolutional layer stored in the configuration register 531. For example, as Figure 6A shown, the secondary processing module 538 reads the two-stage compressed main region 52111 of the input data block (1, 1) of the input feature Figure 1 of the current convolutional layer in the storage segment 521 until the two-stage compressed main region 521M1 of the input data block (1, 1) of the input feature map M; thus, the secondary processing module 538 can read a total of M main regions belonging to different input feature maps. In another embodiment, the secondary processing module 538 can read the two-stage compressed main regions of a part (such as 5) of the input data blocks of each input feature map at a time. Then step S605F is executed.

[0176] In step S605F, the secondary processing module 538 performs secondary decompression on the two-stage compressed main regions of all the read input data blocks and stores them in the cache 532. Specifically, the secondary processing module 538 performs secondary decompression on the two-stage compressed main regions of all the read input data blocks to generate the one-stage compressed main regions of all the input data blocks. Then, the secondary processing module 538 stores the one-stage compressed main regions of all the input data blocks in the cache segment 5321 of the cache 532. For example, the secondary processing module 538 stores the input feature Figure 1The first-level compressed data generated after decompressing the second-level compressed main region 52111 is stored in the main cache segment 532111 of the input feature map cache segment 53211... and so on until the first-level compressed data generated after decompressing the second-level compressed main region 521M1 of the input feature map M is stored in the main cache segment 5321M1 of the input feature map cache segment 5321M. Then, step S607F is executed.

[0177] In step S607F, the convolution operation device 530 determines whether it is necessary to read the first region of the input data block to which the just-read main region belongs. Specifically, in the first embodiment, the secondary processing module 538 reads only the first region of one input data block each time. As Figure 4 shown in the input feature map 410, when processing the input data blocks of the input feature map in the order from left to right and from top to bottom, if the input data block is located in the last row of the input feature map, the judgment result is "no"; if the input data block is not located in the last row of the input feature map, the judgment result is "yes". Similarly, when processing the input data blocks of the input feature map in the order from top to bottom and from left to right, if the input data block is located in the last column of the input feature map, the judgment result is "no"; if the input data block is not located in the last column of the input feature map, the judgment result is "yes". In the second embodiment, the secondary processing module 538 reads the first regions of all input data blocks in the same row (or column) as the read input data block each time. As Figure 4As shown in the input feature map 410, when processing the input data blocks of the input feature map in the order from left to right and from top to bottom, if the input data block to which the read main region (i.e., the horizontal main region) belongs is in the first column, it indicates that the convolution operation device 530 has just started processing a new row of input data blocks. Therefore, it is necessary to read the first region of the input data block (i.e., the lower horizontal row overlapping region). So the judgment result is "yes". However, if the input data block to which the read main region belongs is in the last row, since the input data block in the last row has no first region, there is no need to read the first region. So the judgment result is "no". If the input data block to which the read main region (i.e., the horizontal main region) belongs is neither in the first column nor in the last row, since the first region of the input data block has already been read when processing the input data block in the first column of the same row, there is no need to read it again. So the judgment result is "no". Similarly, when processing the input data blocks of the input feature map in the order from top to bottom and from left to right, if the input data block to which the read main region (i.e., the vertical main region) belongs is in the first row, it indicates that the convolution operation device 530 has just started processing a new column of input data blocks. Therefore, it is also necessary to read the first region of the input data block (i.e., the right vertical column overlapping region). So the judgment result is "yes". However, if the input data block to which the read main region belongs is in the last column, since the input data block in the last column has no first region, there is no need to read the first region. So the judgment result is "no". If the input data block to which the read main region (i.e., the vertical main region) belongs is neither in the first row nor in the last column, since the first region of the input data block has already been read when processing the input data block in the first row of the same column, there is no need to read it again. So the judgment result is "no". In step S607F, if the judgment result is "no", step S613F is executed. If the judgment result is "yes", step S609F is executed. Now, step S609F will be described first.

[0178] In step S609F, the secondary processing module 538 reads the first region of the input data block to which the just-read main region belongs from the memory 520, performs secondary decompression on it, and stores it in the cache 532. Specifically, the secondary processing module 538 reads the first region of the input data block from the storage segment 521 of the memory 520 according to the position of the input data block to which the just-read main region belongs. In the first embodiment, the secondary processing module 538 only reads its own first region of the input data block. For example, as Figure 4 shown, when processing the input data blocks in the order from left to right and from top to bottom, the first region of the input data block (2, 2) of the input feature map 410 is T 2,1 +H 2,2 +T 2,2; when processing the input data blocks in the order from top to bottom and from left to right, the first region of the input data block (2, 2) of the input feature map 410 is T 1,2 +F 2,2 +T 2,2 。In the second embodiment, the secondary processing module 538 reads the first regions of all the input data blocks that are in the same row (or column) as the read input data block. For example, as Figure 4 shown, when processing the input data blocks in the order from left to right and from top to bottom, the first regions of all the input data blocks of the input feature map 410 that are in the same row as the input data block (1, 1) (i.e., the lower horizontal row overlapping region of the input data block) are: H 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 +…。When processing the input data blocks in the order from top to bottom and from left to right, the first regions of all the input data blocks of the input feature map 410 that are in the same column as the input data block (1, 1) (i.e., the right vertical column overlapping region of the input data block) are: F 1,1 +T 1,1 +F 2,1 +T 2,1 +H 3,1 +T 3,1 +…。Then, the secondary processing module 538 performs secondary decompression on the read first regions to generate the first regions that have been compressed at the first level, and stores the first regions that have been compressed at the first level into the sub-cache segment 532113 of the input feature map cache segment 53211 of the cache segment 5321 of the cache 532 respectively, … and so on until the sub-cache segment 5321M3 of the input feature map cache segment 532M1. Then execute S613F.

[0179] Since the memory 520 is located outside the convolution operation module 530, the speed at which the convolution operation module 530 reads the data of the input feature map of the current convolution layer will be affected by the data transmission bandwidth between the memory 520 and the convolution operation module 530. By storing the input feature map data that has been compressed at two levels in the memory 520, the amount of data that needs to be transmitted between the memory 520 and the convolution operation module 530 is reduced, the data transmission efficiency is improved, and thus the efficiency of the convolution operation module 530 performing convolution operations is improved. At the same time, since the data stored in the cache 532 of the convolution operation module 530 is the input feature map data that has been compressed at the first level instead of the uncompressed original data, more input feature map data can be stored in the cache 532, so that the convolution operation module 530 can perform convolution operations on convolution layers with more input feature maps.

[0180] In step S607F, when the convolution operation device 530 determines that the result of judging whether it is necessary to read the result of the first region of the input data block to which the just-read main region belongs is "no", step S613F is executed.

[0181] In step S613F, the primary processing module 534 reads all the primarily compressed main regions from the cache, decompresses them primarily, and stores them in the register 5342. Specifically, the primary processing module 534 reads all the primarily compressed main regions from the main cache segments 532113 to 5321M3 of the input feature map cache segments 53211 to 5321M of the cache segment 5321 of the cache 532, decompresses each primarily compressed main region primarily, and stores them in the sub-register segments 534211 to 53421M of the main register segment 53421 of the register 5342 respectively, and then deletes all the primarily compressed main regions stored in the cache 532. Then step S615F is executed.

[0182] In step S615F, the computing device 500 determines whether it is necessary to read the first region of the input data block to which the just-read main region belongs. The specific determination method is similar to that in step S607F and will not be elaborated here. When the determination result is "no", step S619F is executed. When the determination result is "yes", step S617F is executed. First, step S617F will be described below.

[0183] In step S617F, the primary processing module 534 reads each primarily compressed first region from the cache 532, decompresses it primarily, and stores it in the register 5342. Specifically, the primary processing module 534 reads each primarily compressed first region (532113 - 5321M3) from the secondary cache segments of each input feature map cache segment (53211 - 5321M) of the cache segment 5321 of the cache 532, decompresses each primarily compressed first region primarily, and stores it in the sub-register segments 5342311 to 534231M (or sub-register segments 5342331 to 534233M) of the secondary register segment 53423 of the register 5342, and then releases the storage space occupied by the just-read first region in the cache 532. As Figure 4As shown in the input feature map 410, when processing the input data blocks in the order from left to right and from top to bottom, to generate the input data blocks in the first row, only one first region corresponding to the input data blocks in the first row is required. However, to generate the input data blocks in the second row, in addition to the first region corresponding to the input data blocks in the second row, one first region corresponding to the input data blocks in the first row (i.e., the second region of the input data blocks in the second row) is also required. After generating the input data blocks in the second row, when generating the input data blocks in the third row, the first region corresponding to the input data blocks in the first row is no longer needed. For example, as Figure 4 shown, to generate the input data blocks in the first row of the input feature map 410, only the first region (i.e., the lower horizontal row overlapping region) H corresponding to all the input data blocks in the first row is required 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 … To generate the input data blocks in the second row of the input feature map 410, in addition to the first region (i.e., the lower horizontal row overlapping region) H corresponding to all the input data blocks in the second row 2,1 +T 2,1 +H 2,2 +T 2,2 +H 2,3 +T 2,3 …, the first region (i.e., the second region of all the input data blocks in the second row) H corresponding to all the input data blocks in the first row is also required 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 … After generating the input data blocks in the second row of the input feature map 410, when generating the input data blocks in the third row, the first region H corresponding to all the input data blocks in the first row is no longer needed 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3… Therefore, when generating the input data blocks of all rows, at most two sub-regions (i.e., the first sub-region and the second sub-region) of each input data block in one row of each input feature map in the current convolutional layer need to be saved in the sub-temporary segment 53423 of the register 534 simultaneously. Each time the primary processing module 534 writes a new first sub-region, it is necessary to determine whether the lower horizontal row overlapping regions stored in the sub-sub-temporary segments 5342311 to 534231M and 5342331 to 534233M of the sub-temporary segment 5342 of the register 534 have been used up. If the judgment result is that they have been used up, then use a new lower horizontal row overlapping region to overwrite them. For example, as Figure 4 shown, when generating the first input data block (3, 1) of the 3rd row of the input feature map 410, the first sub-regions H 1,1 +T 1,1 +H 1,2 +T 1,2 +H 1,3 +T 1,3 … of all input data blocks in the 1st row are used up. When processing the input data blocks in the order from top to bottom and from left to right, the processing method is similar to that when processing the input data blocks in the order from left to right and from top to bottom, so it will not be elaborated here. Then step S619F is executed.

[0184] In step S619F, the primary processing module 534 generates input data blocks based on the primary region and sub-regions of the input data blocks stored in the register 5342. Specifically, first, the primary processing module 534 can calculate the starting positions of the first region and the second region of the input data block in the sub-temporary segment 53421 of the register 5342 according to the column number of the input data block to which the primary region of the input data block stored in the register 5342 belongs. Taking Figure 4 the input feature map 410 in it as an example, the starting positions of the first region T 3,1 +H 3,3 +T 3,3 and the second region T 2,2 +H 2,3 +T 2,3 of the input data block (3, 3) in the sub-temporary segment 53421 of the register 5342 are 2*(w-(k-s)) (or 2*(w-(h-s))).

[0185] Then, the primary processing module 534 can obtain the first region and the second region of the input data block from the sub-temporary segment 53423 according to the starting positions of the first region and the second region of the input data block in the sub-temporary segment 53421 of the register 5342.

[0186] Finally, the primary processing module 534 combines the primary area, the first secondary area, and the second secondary area of all input data blocks to generate an input data block. Then, step S621F is executed.

[0187] In step S621F, the primary processing module 534 determines whether the just-generated input data block is the first input data block of the input feature map. If "no", step S625F is executed. If "yes", step S623F is executed. First, step S623F will be described below.

[0188] In step S623F, the primary processing module 534 reads the convolution kernel group from the cache 532 and stores it in the register 5342. Specifically, the primary processing module 534 reads the convolution kernel group (including convolution kernels 1 - M) from the cache segment 5323 of the cache 532, and stores the read convolution kernel group in the sub-registers 534251 - 53425M of the convolution kernel group storage segment 53425 of the register 5342. Then, step S625F is executed.

[0189] In step S625F, the convolution operation module 530 performs a convolution operation on the input data block of each input feature map and the corresponding convolution kernel in the convolution kernel group to generate the corresponding output data block in the output feature map. Specifically, the primary processing module 534 sends all the input data blocks of the input feature maps and the corresponding convolution kernels in the convolution kernel group (one input data block corresponds to one convolution kernel) to the arithmetic unit 536. The arithmetic unit 536 distributes all the received input data blocks and the prefetched corresponding convolution kernels to the idle arithmetic units 5361 - 536Z for convolution operation (for the detailed process of the convolution operation, see the description in the previous text regarding Figure 2 ), and generates the corresponding output data block in the output feature map. The arithmetic unit 536 sends the generated output data block to the data processing module 539. Then, step S627F is executed.

[0190] In step S627F, the convolution operation module 530 determines whether all the output data blocks of the output feature map have been generated. If "no", the convolution operation module 530 will execute steps S603F - S627F again to generate the next output data block of the output feature map. If "yes", step S629F is executed.

[0191] In step S629F, the convolution operation module 530 generates the output feature map. Specifically, after generating the output feature map, the data processing module 539 stores the generated output feature map in the memory 520 after two compressions through the Figure 6C shown processing flow.

[0192] By executing Figure 6F-1 to 6F-2The processing flow shown can generate the next output feature map of the current convolutional layer by reading the next convolutional kernel group in step S601F until all the output feature maps of the current convolutional layer are generated. After all the output feature maps of the current convolutional layer are generated, the convolutional operation module 530 will notify (e.g., in an interrupt manner) the computing device 500. Then the computing device 500 writes the parameters of the next convolutional layer into the configuration register 531 and notifies the convolutional operation module 530 to start the operation of the next convolutional layer until the operation of the entire neural network is completed.

[0193] Figure 7 The processing flow for the computing device 500 to decompress the input data block according to an embodiment of the present invention is shown as follows. Figure 7 As shown, the computing device 500 first reads the input data block (step S701), and then performs primary decompression on the input data block (step S703). First, step S701 is executed.

[0194] In step S701, the primary processing module 534 reads the input data block. For the detailed process, please refer to the description of step S613F above. Then step S703 is executed. Figure 6F-1 to 6F-2 For step S613F above. Then step S703 is executed.

[0195] In step S703, the primary processing module 534 performs primary decompression on the input data block. For the detailed process, please refer to the description of steps S613F - S617F above. As for steps S619F - S627F, they describe generating the input data block from the primary and secondary regions after primary decompression, performing convolutional operations, and generating the output data block, which will not be elaborated here. Figure 6F-1 to 6F-2 For steps S613F - S617F above. As for steps S619F - S627F, they describe generating the input data block from the primary and secondary regions after primary decompression, performing convolutional operations, and generating the output data block, which will not be elaborated here.

[0196] In another embodiment, when the cache space of the cache 532 in the convolutional operation device 530 is relatively abundant, the secondary processing module 538 can read more primary regions of the input data block each time to accelerate its convolutional operation speed.

[0197] Figure 8 The block diagram of the computing device 800 including a convolutional operation module according to another embodiment of the present invention is shown. Different from the computing device 500, the computing device 800 directly stores the output feature map generated after convolutional operation (i.e., the input feature map of the next convolutional layer) in the cache (instead of storing it in the memory), thereby avoiding storing and reading the input feature map of the next convolutional layer in the memory, which can further improve the operation efficiency of the computing device 800. The computing device 800 will be introduced below with reference to FIGS. 9A - 9F - 1 to 9F - 2.

[0198] As Figure 8As shown, the computing device 800 includes a memory 820 and a convolution operation module 830. The memory 820 is coupled to the convolution operation module 830. The convolution operation module 830 includes a configuration register 531, a cache 832, a data processing module 839, a primary processing module 534, a secondary processing module 838, and an arithmetic unit 536. The data processing module 839 is coupled to the secondary processing module 838 and the arithmetic unit 536. The secondary processing module 838 is coupled to the cache 832 and the data processing module 839. The primary processing module 534 is coupled to the cache 832 and the arithmetic unit 536. The configuration register 531, the primary processing module 534, and the arithmetic unit 536 in the convolution operation module 830 are the same as the configuration register 531, the primary processing module 534, and the arithmetic unit 536 in the convolution operation device 500 respectively, and will not be elaborated here. The following introduces the cache 832, the secondary processing module 838, and the data processing module 839.

[0199] The cache 832 includes cache segments 5321, 5323, and 8322. The cache segments 5321 and 5323 are the same as the cache segments 5321 and 5323 in Figure 5 and will not be elaborated here. The cache segment 8322 is used to store the input feature map data of the next convolution layer (which will be detailed later). The data processing module 839 includes a segmentation module 535 and a compression module 837. The compression module 837 is coupled to the segmentation module 535. The segmentation module 535 is the same as the segmentation module 535 in the data processing module 539 in Figure 5 and will not be elaborated here. As described above, after the data processing module 839 receives the output feature map generated by the arithmetic unit 536 (i.e., the input feature map of the next convolution layer), the segmentation module 535 segments the output feature map into output data blocks (i.e., the input data blocks of the next convolution layer), and then sends them to the compression module 837. The compression module 837 performs primary compression on the received output data blocks and then sends them to the secondary processing module 838, and then the secondary processing module 838 stores the output data blocks after primary compression into the cache segment 8322 of the cache 832. Different from the computing device 500, the data processing module 839 directly stores the output data blocks after primary compression into the cache 832 through the secondary processing module 838 (instead of first storing them into the memory 820 and then having the secondary processing module 838 read them back from the memory), thus reducing the data transmission between the convolution operation module 830 and the memory 820. If the output feature map generated by the arithmetic unit 536 is the output feature map of the last convolution layer, the data processing module 839 will directly store the received output feature map into the memory 820.

[0200] Since the input feature map of the first convolutional layer (stored in the memory 820) is the original input data for the convolutional operation, it needs to be compressed at the first level and then stored in the cache 832 before the convolutional operation is performed by the computing device 800. Specifically, the computing device 800 reads the input feature map of the first convolutional layer from the storage segment 821 of the memory 820 as shown in Figure 9A and then sends it to the data processing module 839. The data processing module 839 then stores the received input feature map of the first convolutional layer in the cache 832 after performing segmentation and compression processing through the segmentation module 535 and the compression module 837. The specific segmentation and compression processes have been described before and will not be elaborated here. In one embodiment, the convolutional operation module 830 also provides a decompression / compression interface externally; through this decompression / compression interface, a module located outside the convolutional operation module 830 can call the data processing module 839 to perform decompression / compression operations; at this time, the data processing module 839 is simply called.

[0201] Figure 9A FIG. is a schematic diagram of the data stored in the memory 820 of the computing device 800 according to an embodiment of the present invention. As shown in Figure 9A , the memory 820 includes storage segments 821, 823, 525, and 527. The storage segments 525 and 527 are the same as the storage segments 525 and 527 of the memory 520 and will not be elaborated here. The storage segment 821 is used to store the set of input feature maps for the convolutional operation (as described in the previous paragraph, that is, the set of input feature maps of the first convolutional layer), and the storage segment 823 is used to store the set of output feature maps for the convolutional operation (the set of output feature maps of the last convolutional layer).

[0202] Figure 9B FIG. is a more detailed block diagram of the computing device 800 according to an embodiment of the present invention. As shown in Figure 9B , the configuration register 531, the arithmetic unit 536, the first-level processing module 534, the cache segment 5321, the cache segment 5323, and Figure 6BThe configuration register 531, arithmetic unit 536, primary processing module 534, cache segment 5321, and cache segment 5323 are the same as those described above, and will not be elaborated here. The cache segment 8322 is used to store the data of the input feature map of the next convolutional layer, and its storage structure is exactly the same as that of the cache segment 5321. The difference is that the cache segment 5321 is used to store the data of the input feature map of the current convolutional layer, while the cache segment 8322 is used to store the data of the input feature map of the next convolutional layer. In one embodiment, the cache segments 5321 and 8322 can be alternately used to store the data of the input feature maps of the current convolutional layer and the next convolutional layer. For example, during the convolution operation on the input data block of the Nth layer, the cache segment 5321 is used to store the data of the input feature map of the current convolutional layer (i.e., the Nth convolutional layer), and the cache segment 8322 is used to store the data of the input feature map of the next convolutional layer (i.e., the (N + 1)th convolutional layer). During the convolution operation on the input data block of the (N + 1)th layer, the cache segment 8322 is used to store the data of the input feature map of the current convolutional layer (i.e., the (N + 1)th convolutional layer), and the cache segment 5321 is used to store the data of the input feature map of the next convolutional layer (i.e., the (N + 2)th convolutional layer), and so on.

[0203] Figure 9C FIG. is a processing flow for writing the input feature map of the Nth convolutional layer into the cache after primary compression according to an embodiment of the present invention. As Figure 9C shown, the data processing module 839 first generates an input data block (step S901C), then performs primary compression on the input data block (step 903C), and finally stores the input data block after primary compression into the cache 832 (step S907C). Figure 9C Steps S901C and S903C in Figure 6C are the same as steps S601C and S603C in

[0204] and will not be elaborated here. The following describes step S907C.

[0204] In step S907C, the secondary processing module 838 stores the input data block after primary compression into the cache 832. Specifically, the secondary processing module 838 stores the primary region and secondary region of each input data block of the input feature map after primary compression into the cache segment 8322 of the cache 832 (for example, stores the input feature map of the Nth convolutional layer into the cache segment 8322) or the cache segment 5321 (for example, stores the input feature map of the (N + 1)th convolutional layer into the cache segment 5321, that is, stores the output feature map of the Nth convolutional layer into the cache segment 5321).

[0205] After receiving the notification to start the convolution operation, the computing device 800 will use Figure 9D or Figure 9EThe processing flow in (to be described in detail later) performs a convolution operation on the input feature map set of the first convolutional layer and each convolutional kernel group to generate an output feature map corresponding to each convolutional kernel group. As Figure 9D shown, the processing flow for the computing device 800 to perform a convolution operation on the input feature map set and a convolutional kernel group to generate an output feature map is as follows: each of the multiple input data blocks is divided into multiple non-overlapping regions, where there is an overlapping region between any two adjacent input data blocks (step S903D); the multiple non-overlapping regions of each input data block are stored in their respective corresponding non-overlapping storage spaces in the cache (S905D); each input data block is generated according to the regions corresponding to each input data block stored in the non-overlapping storage spaces (S907D); a convolution operation is performed on the generated multiple input data blocks to generate the output feature map (S909D). Figure 9D Steps S903D, S907D, and S909D in Figure 6D are the same as steps S903D, S907D, and S909D in

[0206] and will not be elaborated here. Next, step S905D is introduced.

[0206] In step S905D, the computing device 800 stores the multiple non-overlapping regions of each input data block in their respective corresponding non-overlapping storage spaces in the cache. Specifically, the secondary processing module 838 of the computing device 800 performs primary compression on the multiple non-overlapping regions of the multiple input data blocks generated in step S903D and stores them in the cache segment 8322 or 5321 of the cache 832.

[0207] Figure 9E is the processing flow for the computing device 800 to generate an output feature map according to another embodiment of the present invention. As Figure 9E shown, the processing flow for the computing device 800 to generate an output feature map is as follows: perform a primary decompression operation on the main region and at least one secondary region of each input data block (step S905E); generate each input data block by using the main region and the secondary region after the primary decompression operation of each input data block (S907E); perform a convolution operation on each input data block to generate the output feature map (S909E). Figure 9E Steps S907E and S909E in Figure 6E are the same as steps S607E and S609E in

[0208] In step S905E, the computing device 800 performs a first-level decompression operation on the main region and at least one secondary region of each of the input data blocks. Specifically, the computing device 500 reads the first-level compressed main region and secondary region of the input data block from the cache 532, performs a first-level decompression operation on the read first-level compressed main region and secondary region, and stores the result in the register 5342. For a more detailed process, refer to the description of step S913F of Figure 9F-1 to 9F-2 below.

[0209] Figure 9F-1 to 9F-2 FIG. shows a more detailed processing flow for the computing device 800 to generate an output feature map according to an embodiment of the present invention. As shown, Figure 9F-1 to 9F-2 describes the processing flow in which the computing device 800 performs a convolution operation on an input feature map set and a convolution kernel group to generate an output feature map during the convolution operation. When the space in the cache 832 is large enough, during the convolution operation, the computing device 800 directly stores the output feature map generated by each convolution layer (excluding the last convolution layer, the output feature map of the last convolution layer will be directly stored in the memory 820) into the cache 832 after segmentation and first-level compression, instead of sending the output feature map generated by each convolution layer to the memory 820 for storage and then loading it from the memory 820 to the convolution operation module 830 for processing. This can reduce the data transmission between the convolution operation module 830 and the memory 820, and thus improve the efficiency of the entire system when performing the convolution operation.

[0210] Figure 9F-1 to 9F-2 It includes steps S901F, S913F, S915F, S917F, S919F, S921F, S923F, S925F, S927F, and S929F. Among them, steps S901F, S913F, S915F, S917F, S919F, S921F, S923F, S925F, and S929F are the same as steps S601F, S613F, S615F, S617F, S619F, S621F, S623F, S625F, and S629F in Figure 6F-1 to 6F-2 and will not be elaborated here. Different from Figure 6F-1 to 6F-2 , in step S927F, when the convolution operation module 830 determines whether all output data blocks of the output feature map have been generated, if the determination result is no, step S913F is executed.

[0211] With the data decompression method, data compression method, and convolution operation device described in the present application, by compressing and storing the input data blocks, more input data blocks can be cached in the convolution operation device, thereby reducing the pause times of the convolution operation module, and thus improving the operation efficiency of the convolution operation module.

[0212] Although the present application has been disclosed above by way of examples, it is not intended to limit the present application. Any person skilled in the art can make various modifications and refinements without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application shall be subject to the scope defined by the appended claims.

Claims

1. A data decompression method, used in a convolution operation device, to decompress an input data block, the data decompression method comprising: Reading the input data block; and Performing a first-level decompression on the input data block, in, The first-level decompression is used to decompress the input data block whose data format is a first-level compression algorithm format; The first-level compression algorithm format includes: a mask field, used to identify the position of non-zero elements in the input data block and the number of elements in the input data block, wherein the length of the mask field is not less than the number of elements in the input data block, Among them, each element in the area of ​​the input data block corresponds to a bit in the mask field, the last bit with a value of 1 in the mask field corresponds to the last element in the area of ​​the input data block, a bit with a value of 0 in the mask field corresponds to each element with a value of 0 in the area of ​​the input data block except the last element, and a bit with a value of 1 in the mask field corresponds to each element with a value not equal to 0 in the area of ​​the input data block except the last element.

2. The data decompression method according to claim 1, wherein: The first-level compression algorithm format also includes: a target data field, used to store the last element in the region of the input data block and all non-zero elements in the region of the input data block except the last element; and The length field is used to indicate the number of elements in the target data field.

3. The data decompression method according to claim 1, wherein: Each element in the region of the input data block whose value is not 0 corresponds to a bit in the mask field whose value is 1.

4. The data decompression method according to claim 3, wherein: The number of bits in the mask field is greater than the number of elements in the region of the input data block.

5. The data decompression method according to claim 1, wherein: The first-level compression algorithm format also includes: A target data field is used to store all non-zero elements in the region of the input data block; and The length field is used to indicate the number of elements in the target data field.

6. The data decompression method according to claim 5, wherein: The maximum number of elements that the length field can indicate is greater than the maximum value of the length field, and the number of elements in the target data field is greater than the number of all non-zero elements in the region of the input data block.

7. A data compression method, used in a convolution operation device, to compress an input data block, the data compression method comprising: generating the input data block; and Performing a first level compression on the input data block, in, The first-level compression compresses the input data block into data in a first-level compression algorithm format; The first-level compression algorithm format includes: a mask field, used to identify the position of non-zero elements in the input data block and the number of elements in the input data block, wherein the length of the mask field is not less than the number of elements in the input data block, Among them, each element in the area of ​​the input data block corresponds to a bit in the mask field, the last bit with a value of 1 in the mask field corresponds to the last element in the area of ​​the input data block, a bit with a value of 0 in the mask field corresponds to each element with a value of 0 in the area of ​​the input data block except the last element, and a bit with a value of 1 in the mask field corresponds to each element with a value not equal to 0 in the area of ​​the input data block except the last element.

8. The data compression method according to claim 7, wherein: The first-level compression algorithm format also includes: a target data field, used to store the last element in the region of the input data block and all non-zero elements in the region of the input data block except the last element; and The length field is used to indicate the number of elements in the target data field.

9. The data compression method according to claim 7, wherein: Each element in the region of the input data block whose value is not 0 corresponds to a bit in the mask field whose value is 1.

10. The data compression method according to claim 9, wherein: The number of bits in the mask field is greater than the number of elements in the region of the input data block.

11. A convolution operation device for decompressing an input data block, the convolution operation device comprising: A cache, for storing the input data block; and A primary processing module, reading the input data block from the cache and performing primary decompression on the input data block; in, The first-level decompression is used to decompress the input data block whose data format is a first-level compression algorithm format; The first-level compression algorithm format includes: a mask field, used to identify the position of non-zero elements in the input data block and the number of elements in the input data block, wherein the length of the mask field is not less than the number of elements in the input data block, Among them, each element in the area of ​​the input data block corresponds to a bit in the mask field, the last bit with a value of 1 in the mask field corresponds to the last element in the area of ​​the input data block, a bit with a value of 0 in the mask field corresponds to each element with a value of 0 in the area of ​​the input data block except the last element, and a bit with a value of 1 in the mask field corresponds to each element with a value not equal to 0 in the area of ​​the input data block except the last element.

12. The convolution operation device according to claim 11, wherein: The first-level compression algorithm format also includes: a target data field, used to store the last element in the region of the input data block and all non-zero elements in the region of the input data block except the last element; and The length field is used to indicate the number of elements in the target data field.

13. The convolution operation device according to claim 11, wherein: Each element in the region of the input data block whose value is not 0 corresponds to a bit in the mask field whose value is 1.

14. The convolution operation device according to claim 13, wherein: The number of bits in the mask field is greater than the number of elements in the region of the input data block.

15. The convolution operation device according to claim 11, wherein: The first-level compression algorithm format also includes: A target data field is used to store all non-zero elements in the region of the input data block; and The length field is used to indicate the number of elements in the target data field.

16. The convolution operation device according to claim 15, wherein: The maximum number of elements that the length field can indicate is greater than the maximum value of the length field, and the number of elements in the target data field is greater than the number of all non-zero elements in the region of the input data block.

17. A convolution operation device for compressing an input data block, the convolution operation device comprising: A cache, for storing the input data block; and A data processing module performs primary compression on the input data block and stores the primary compressed input data block into the cache; in, The first-level compression compresses the input data block into data in a first-level compression algorithm format; The first-level compression algorithm format includes: a mask field, used to identify the position of non-zero elements in the input data block and the number of elements in the input data block, wherein the length of the mask field is not less than the number of elements in the input data block, Among them, each element in the area of ​​the input data block corresponds to a bit in the mask field, the last bit with a value of 1 in the mask field corresponds to the last element in the area of ​​the input data block, a bit with a value of 0 in the mask field corresponds to each element with a value of 0 in the area of ​​the input data block except the last element, and a bit with a value of 1 in the mask field corresponds to each element with a value not equal to 0 in the area of ​​the input data block except the last element.

18. The convolution operation device according to claim 17, wherein: The first-level compression algorithm format also includes: a target data field, used to store the last element in the region of the input data block and all non-zero elements in the region of the input data block except the last element; and The length field is used to indicate the number of elements in the target data field.

19. The convolution operation device according to claim 17, wherein: Each element in the region of the input data block whose value is not 0 corresponds to a bit in the mask field whose value is 1.

20. The convolution operation device according to claim 19, wherein: The number of bits in the mask field is greater than the number of elements in the region of the input data block.

Citation Information

Patent Citations

  • Neural network processor using compression and decompression of activation data to reduce memory bandwidth utilization

    CN110520909A