Method for reducing memory bandwidth
By splitting the multi-layer convolution operation into multi-level operation and using cache memory temporary storage results, the problem of high memory bandwidth demand in the prior art is solved, and the memory bandwidth reduction and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202211296441.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-11-29
AI Technical Summary
In the prior art, multi-layer convolutional operation of neural network models requires a large amount of memory bandwidth, resulting in frequent read and write of dynamic random access memory and high bandwidth requirements.
The multi-layer convolution operation is split into multi-level operations, and the results of each-level operation are temporarily stored in the cache memory, reducing the number of accesses to the dynamic random access memory, and optimizing data transmission through the memory management circuit.
It reduces the bandwidth requirement of memory, reduces the read and write frequency of memory, and improves the efficiency of memory usage.
Smart Images

Figure CN115660055B_ABST
Abstract
Description
[0001] This application is a divisional application of the application submitted to the Chinese Patent Office on November 29, 2021, with the application number 202111433001.1 and the invention title "Intelligent Processor Device and Method for Reducing Memory Bandwidth", the entire content of which is incorporated herein by reference. Technical Field
[0002] This application relates to the technical field of intelligent processors, and specifically relates to a method for reducing memory bandwidth. Background Art
[0003] Existing neural network models usually include multiple layers of convolutional operations performed in sequence. As Figure 1 shown, in the prior art, the input feature map data in the dynamic random access memory is split into multiple tile data (drawn with dashed lines). In the first layer of convolutional operation, the processor sequentially obtains and processes these tile data to generate new tile data, and writes the new tile data back to the dynamic random access memory in sequence. Then, when performing the second layer of convolutional operation, the processor reads out and processes the multiple tile data obtained in the first layer of convolutional operation from the dynamic random access memory in sequence to generate new tile data, and writes the new tile data back to the dynamic random access memory in sequence, and so on until all layers of convolutional operations are completed. In other words, in the prior art, the data output by each layer of convolutional operation is used as the input data for the next layer, so the dynamic random access memory needs to be repeatedly read and written. In this way, in the prior art, the dynamic random access memory needs to have a large memory bandwidth to be sufficient to perform multiple layers of convolutional operations. Summary of the Invention
[0004] An embodiment of this application provides a method for reducing memory bandwidth, aiming to reduce the memory bandwidth requirement.
[0005] In some embodiments, the method for reducing memory bandwidth can be applied to an intelligent processor device that executes a convolutional neural network model. The intelligent processor device includes a first memory, a memory management circuit, a second memory, and a convolutional operation circuit. The method for reducing memory bandwidth includes the following operations: determining the data size of a block of data stored in the first memory when the convolutional operation circuit performs a convolutional operation according to the capacity of the first memory. The memory management circuit transfers the block of data from a dynamic random access memory to the first memory, and the convolutional operation circuit obtains the block of data from the first memory and sequentially performs multiple levels of operations corresponding to the convolutional operation on the block of data to sequentially generate multiple output feature map data; determining the number of levels of the multiple levels of operations and the required data amount of at least a second part of each of the remaining data in the multiple output feature map data according to the capacity of the second memory and the data amount of a first part of data of the last output feature map data in the multiple output feature map data. During the execution of the multiple levels of operations, the memory management circuit stores the first part of data and the at least a second part of data in the second memory; and generating a predetermined compilation file based on the data size of the block of data, the number of levels of the multiple levels of operations, the data amount of the first part of data, and the data amount of the at least a second part of data. The memory management circuit accesses the dynamic random access memory, the first memory, and the second memory based on the predetermined compilation file.
[0006] Regarding the features, implementation, and effects of the present application, the following provides a detailed description of the preferred embodiments in conjunction with the drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] To more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0008] Figure 1 Conceptual schematic diagram of performing a convolutional operation for the prior art;
[0009] Figure 2 Schematic diagram of an artificial intelligence system drawn according to some embodiments of the present application;
[0010] Figure 3A Conceptual schematic diagram of the basic concept of a convolutional operation drawn according to some embodiments of the present application;
[0011] Figure 3B For the amount drawn according to some embodiments of the present application Figure 2Conceptual schematic diagram of a convolutional operation performed by the intelligent processor device in
[0012] Figure 3C Drawn according to some embodiments of the present application Figure 2 Schematic diagram of the data transfer process of the intelligent processor device in
[0013] Figure 4 Flowchart of a method for reducing memory bandwidth drawn according to some embodiments of the present application
[0014] Figure 5A Drawn according to some embodiments of the present application Figure 4 Conceptual schematic diagram of an operation in ; and
[0015] Figure 5B Drawn according to some embodiments of the present application Figure 4 Detailed step flowchart of an operation in Detailed implementation manners
[0016] All terms used herein have their ordinary meanings. The definitions of the above terms in commonly used dictionaries are only examples of the use of any term discussed herein and should not limit the scope and meaning of the present application. Similarly, the present application is not limited only to the various embodiments shown in this specification.
[0017] Regarding the use of "coupled" or "connected" herein, it may refer to two or more components making direct physical or electrical contact with each other, or making indirect physical or electrical contact with each other, and may also refer to two or more components operating or acting on each other. As used herein, the term "circuit" may be a device that processes signals by connecting at least one transistor and / or at least one active or passive component in a certain manner.
[0018] In some embodiments, an intelligent processor device (such as the intelligent processor device 230 of Figure 2 ) may split the multi-layer convolutional operations in the convolutional neural network model into multiple levels of operations, and temporarily store the calculation results generated in the multiple levels of operations in a cache memory (such as the memory 233 of Figure 2 ), and write the finally generated data back to a dynamic random access memory (such as the memory 220 of Figure 2 ) after completing all levels of operations. In this way, the bandwidth requirement of the dynamic random access memory can be reduced.
[0019] Figure 2A schematic diagram of an artificial intelligence system 200 drawn according to some embodiments of the present application. The artificial intelligence system 200 includes a processor 210, a memory 220, and an intelligence processor (or intelligence processing unit) device 230. The artificial intelligence system 200 can be used to execute a neural network model (such as, but not limited to, a convolutional neural network model) to process various types of data (such as, but not limited to, image data).
[0020] The memory 220 can store the input data DI to be processed and the output data DO processed by the intelligence processor device 230. In some embodiments, the memory 220 can be a dynamic random access memory. The intelligence processor device 230 can read the input data DI from the memory 220 based on the control of the processor 210 and perform a convolution operation on the input data DI to generate the output data DO.
[0021] Specifically, the intelligence processor device 230 includes a memory management circuit 231, a memory 232, a memory 233, and a convolution operation circuit 234. The memory management circuit 231 is coupled to the memory 232, the memory 233, and the memory 220. In some embodiments, the memory management circuit 231 can be implemented by (but not limited to) circuits such as a memory management unit and a direct memory access controller. The memory management circuit 231 can read the input data DI from the memory 220 to the memory 232 based on the control of the processor 210 and / or the convolution operation circuit 234. The convolution operation circuit 234 can read the memory 232 to obtain the input data DI and perform a convolution operation on the input data DI to generate the output data DO. After the convolution operation circuit 234 generates the output data DO, the memory management circuit 231 can transfer the output data DO to the memory 220 to store the output data DO.
[0022] In some embodiments, the processor 210 may send an instruction CMD based on a predetermined compilation file (not shown), and the intelligent processor device 230 may read the input data DI from the memory 220 according to the instruction CMD and perform a convolution operation on the input data DI to generate the output data DO. The intelligent processor device 230 may further split the multi-layer convolution operation of the convolutional neural network model into multiple-level operations according to the instruction CMD. The memory management circuit 231 may temporarily store the output results (such as the output feature map data described later) generated by the convolution operation circuit 234 in each level of operation in the memory 233, and the convolution operation circuit 234 may access the memory 233 via the memory management circuit 231 during the operation process to use these temporarily stored data to complete each level of operation, and write the operation result of the last level (equivalent to the output data DO) back to the memory 220 after the operation is completed. In this way, the access times of the intelligent processor device 230 to the memory 220 can be reduced, and thus the memory bandwidth required in the artificial intelligence system 200 can be reduced. The operations here will be described later with reference to Figures 3A to 3C Description.
[0023] In some embodiments, both the memory 232 and the memory 233 are static random access memories. For example, the memory 232 may be a second-level (L2) cache memory, and the memory 233 may be a third-level (L3) cache memory. In some embodiments, the memory 232 is a two-dimensional memory, and the data width of the memory 232 is equal to the data width of the convolution operation circuit 234 and / or the memory management circuit 231. For example, the memory 232 may have 32 channels (slots), the data width of each channel is 256 bits and the depth is 512, but the present application is not limited to the above values. The memory 232 is directly connected to the convolution operation circuit 234 for the convolution operation circuit 234 to directly access. In contrast, the memory 233 is a one-dimensional memory, and its data width is different from the data width of the convolution operation circuit 234 and / or the memory management circuit 231, and the convolution operation circuit 234 can access the memory 233 via the memory management circuit 231 to temporarily store the results generated in each level of operation.
[0024] To illustrate the related operations of the intelligent processor device 230, the basic concept of the convolution operation and the multi-level operation and data transfer process of the intelligent processor device 230 will be described in sequence below.
[0025] Figure 3A It is a schematic diagram of the basic concept of the convolution operation drawn according to some embodiments of the present application. As Figure 3AAs shown, the width and height of the input feature map data 310 are w1 and h1 respectively, and the width and height of the output feature map data 320 are w2 and h2 respectively. In this example, the input feature map data 310 includes four tile data 310[1] to 310[4] (drawn in different line styles), and the output feature map data 320 includes four tile data 320[1] to 320[4] (drawn with different shadings), where the width and height of each of the tile data 310[1] to 310[4] are tile_w1 and tile_h1 respectively, and the width and height of each of the tile data 320[1] to 320[4] are tile_w2 and tile_h2 respectively. During the execution of the convolution operation, the memory management circuit 231 can transfer one tile data in the input feature map data 310 to the memory 232, and the convolution operation circuit 234 can read the tile data from the memory 232 and perform an operation on the tile data using the convolution kernel 315 to generate a corresponding tile data in the output feature map data 320. For example, the memory management circuit 231 can transfer the tile data 310[1] to the memory 232, and the convolution operation circuit 234 can read the tile data 310[1] from the memory 232 and perform an operation on the tile data 310[1] using the convolution kernel 315 to generate the tile data 320[1].
[0026] In some embodiments, the convolution operation circuit 234 can split the multi-layer convolution operation into multiple-stage operations, and the operation mode of each stage is similar to Figure 3A the operation concept shown. The output feature map data generated by the current-stage operation will be used as the input feature map data for the next-stage operation, and the output feature map data of the next-stage operation will be used as the input feature map data for the next-next-stage operation. However, in actual operations, for the same data, the data size of the tile data in the output feature map data of the previous stage is usually different from the data size of the tile data in the input feature map data of the next stage. For example, as Figure 3AAs shown, the data size (e.g., width tile_w1 and height tile_h1) of the block data (e.g., block data 310[1]) in the input feature map data 310 is larger than the data size (e.g., width tile_w2 and height tile_h2) of the block data (e.g., block data 320[1]) of the output feature map data 320. Therefore, the data amount of the block data in the output feature map data of the previous level must be sufficient to be used as at least one block data of the input feature map data in the next level of operation. By setting up the memory 233, after completing one operation in each level of operation, the convolution operation circuit 234 can transmit a block data corresponding to the output feature map data of the current level to the memory 233 via the memory management circuit 231. When the amount of block data stored in the memory 233 accumulates to a sufficient amount, the convolution operation circuit 234 can access the memory 233 via the memory management circuit 231 to obtain the multiple block data, and use the multiple block data as at least one block data of the input feature map data of the next level to perform the convolution operation of the next level to generate at least one block data of the output feature map data of the next level.
[0027] Figure 3B According to some embodiments of the present application Figure 2 A conceptual diagram of the intelligent processor device 230 performing a convolution operation. Figure 1 As shown, in the prior art, the operation result generated by each layer of convolution operation will first be written back to the dynamic random access memory, and then read from the dynamic random access memory to perform the calculation of the next row of data or the next layer of convolution. In this way, the dynamic random access memory needs to have a sufficiently large read and write bandwidth. Compared with the prior art, in some embodiments of the present application, the convolution operation circuit 234 can split the multi-layer convolution operation into multiple levels of operation, each level of operation only completes part of the convolution layer operation, and temporarily stores the operation results generated level by level in the memory 233 (rather than directly storing them back to the memory 220), and reads the operation results of the previous level from the memory 233 when performing the next level of operation. Similarly, after completing all levels of operation, the convolution operation circuit 234 can generate output data DO, and store the output data in the memory 220 via the memory management circuit 231. In this way, the usage bandwidth of the memory 220 can be reduced.
[0028] exist Figure 3BIn [the figure], the block data corresponding to the output feature map data of each level of operation is drawn as a dotted square. The convolution operation circuit 234 performs a first-level operation on the input data DI to generate output feature map data 330-1 corresponding to the first level. The memory management circuit 231 can sequentially store multiple block data in the output feature map data 330-1 in the memory 233. When the data volume of the multiple block data of the output feature map data 330-1 stored in the memory 233 meets a default value (for example, but not limited to, the block data accumulated to one row), the convolution operation circuit 234 reads out the multiple block data from the memory 233 via the memory management circuit 231, and uses the multiple block data (corresponding to the output feature map data 330-1) as at least one block data of the input feature map data for the second-level operation, and performs a second-level operation on the multiple block data to generate block data corresponding to the output feature map data 330-2 of the second level. Among them, the aforementioned preset value is the data volume sufficient for the convolution operation circuit 234 to generate at least one block data of the output feature map data 330-2 (for example, but not limited to, the block data of one row). By analogy, the convolution operation circuit 234 can sequentially perform multiple levels of operations, and sequentially generate multiple output feature map data 330-3, 330-4,..., 330-n, where the output feature map data 330-n of the last level of operation is equivalent to the output data DO corresponding to the input data DI passing through n convolutional layers.
[0029] Figure 3C Schematic diagram of the data transfer process of the intelligent processor device 230 in accordance with some embodiments of the present application. In this example, Figure 2 Figure 3B The multi-layer convolution operation shown can be further split into n-level operations. In the first-level operation, the memory management circuit 231 can read at least one block of data of the input data DI (which is the input feature map data of the first-level operation) from the memory 220 (step S3-11), and store the at least one block of data in the memory 232 (step S3-12). The convolution operation circuit 234 can obtain the at least one block of data from the memory 232, and perform a convolution operation on the at least one block of data to generate at least one block of data in the output feature map data 330-1 (step S3-13). The convolution operation circuit 234 can store at least one block of data in the output feature map data 330-1 in the memory 232 (step S3-14). The memory management circuit 231 can transfer at least one block of data in the output feature map data 330-1 stored in the memory 232 to the memory 233 (steps S3-15 and S3-16). Steps S3-11 to S3-16 are repeatedly executed until the data volume of at least one block of data in the output feature map data 330-1 stored in the memory 233 meets a first default value. The memory management circuit 231 can read the at least one block of data from the memory 233 to enter the second-level operation (step S3-21), and transfer the at least one block of data in the output feature map data 330-1 (which is equivalent to the input feature map data of the second-level operation) to the memory 232 (step S3-22). Among them, the first preset value is the data volume sufficient for the convolution operation circuit 234 to generate at least one block of data in the output feature map data 330-2 (for example, but not limited to, one row of block data).
[0030] Similarly, in the second-level operation, the convolution operation circuit 234 can obtain the at least one block of data from the memory 232, and perform an operation on the at least one block of data to generate at least one block of data in the output feature map data 330-2 (step S3-23). The convolution operation circuit 234 can store at least one block of data in the output feature map data 330-2 in the memory 232 (step S3-24), and the memory management circuit 231 can transfer at least one block of data in the output feature map data 330-2 stored in the memory 232 to the memory 233 (steps S3-25 and S3-26). Steps S3-21 to S3-26 are repeated until the data volume of at least one block of data in the output feature map data 330-2 stored in the memory 233 meets a second default value. The memory management circuit 231 can read the at least one block of data from the memory 233 to enter the third-level operation (not shown), where the second preset value is the data volume sufficient for the convolution operation circuit 234 to generate at least one block of data in the output feature map data 330-3.
[0031] By analogy, when the data volume of at least one block of data in the output feature map data 330-(n-1) stored in the memory 233 meets a specific default value, the memory management circuit 231 can read out the at least one block of data from the memory 233 to enter the nth-level operation (step S3-n1), and store the at least one block of data in the output feature map data 330-(n-1) (which is equivalent to the input feature map data of the nth-level operation) in the memory 232 (step S3-n2), where the specific preset value is the data volume sufficient for the convolution operation circuit 234 to generate at least one block of data in the output feature map data 330-n. In the nth-level operation, the convolution operation circuit 234 can obtain the at least one block of data from the memory 232, and perform a convolution operation on the at least one block of data to generate at least one block of data in the output feature map data 330-n (step S3-n3). The convolution operation circuit 234 can store at least one block of data in the output feature map data 330-(n) in the memory 232 (step S3-n4), and the memory management circuit 231 can transfer at least one block of data in the output feature map data 330-n stored in the memory 232 to the memory 220 (steps S3-n5 and S3-n6).
[0032] In other words, when performing the first-level operation, the block data of the input feature map data is read out from the memory 220. The block data in the input (or output) feature map data generated in the intermediate levels of operations is temporarily stored in the memory 233. When performing the last level (i.e., the nth level) of operation, the final output feature map data (equivalent to the output data DO) is stored in the memory 220. By repeatedly executing the above steps, the convolution operation circuit 234 can complete the operation on all the block data temporarily stored in the memory 233.
[0033] Figure 4 FIG. 400 is a flowchart of a method for reducing memory bandwidth according to some embodiments of the present application. The memory bandwidth reduction method 400 can be applied to various systems or devices (such as, but not limited to, Figure 2 the artificial intelligence system 200) that execute an artificial neural network model to reduce the usage bandwidth of the memory in the system.
[0034] In step S410, according to the capacity of the first memory (such as Figure 2 the memory 232), determine the data size of a block of data stored in the first memory when the convolution operation circuit (such as Figure 2 the convolution operation circuit 234) performs a convolution operation, where the memory management circuit (such as Figure 2 the memory management circuit 231) reads from a dynamic random access memory (such as Figure 2The memory 220) transfers the block data to the first memory, and the convolution operation circuit obtains the block data from the first memory and sequentially performs multiple-level operations corresponding to the convolution operation on the block data to generate multiple output feature map data (for example, Figure 3B Multiple output feature map data 330-1 to 330-n).
[0035] In some embodiments, step S410 can be used to determine the data size of the block data read into the memory 232. As Figure 3C shown, in each level of operation, the block data of the input feature map data (equivalent to the output feature map data generated in the previous level of operation) and the block data of the output feature map data are stored in the memory 232. As Figure 3A shown, the data size of the block data (for example, block data 310[1]) in the input feature map data 310 is different from the data size of the block data (for example, block data 320[1]) in the output feature map data 320. Under this condition, if the data size of the block data read into the memory 232 is smaller, it is necessary to repeatedly access the memory 220 (and / or the memory 232) multiple times to retrieve the complete input feature map data 310. In this way, the bandwidth requirements of the memory 220 and / or the memory 232 will become larger. Therefore, in order to reduce the aforementioned number of reads, the data size of the block data read into the memory 232 can be set as large as possible on the premise of meeting the capacity of the memory 232.
[0036] Specifically, taking Figure 3A as an example, if the capacity of the memory 232 is X, the total data volume of the block data in the input feature map data 310 and the block data in the output feature map data 320 stored in the memory 232 cannot exceed the capacity X of the memory 232, which can be expressed by the following formula (1):
[0037] tile_w1×tile_h1×c1+tile_w2×tile_h2×c2<X…(1)
[0038] where the width tile_w1 and the height tile_h1 are the data sizes of the block data in the input feature map data 310, the width tile_w2 and the height tile_h2 are the data sizes of the block data in the output feature map data 320, c1 is the number of channels corresponding to the input feature map data 310, and c2 is the number of channels corresponding to the output feature map data 320.
[0039] Furthermore, in the mathematical concept of convolution operation, the width tile_w1 and height tile_h1 of the block data in the input feature map data 310, and the width tile_w2 and height tile_h2 of the block data in the output feature map data 320 satisfy the following equations (2) and (3):
[0040] (tile_w1 - f_w) / stride_w + 1 = tile_w2…(2)
[0041] (tile_h1 - f_h) / stride_h + 1 = tile_h2…(3)
[0042] Where f_w and f_h are the width and height of the convolution kernel 315 respectively, stride_w is the width stride of each movement of the convolution kernel 315 on the input feature map data 310, and stride_h is the height stride of each movement of the convolution kernel 315 on the input feature map data 310.
[0043] In addition, since there is overlapping data between multiple block data 310[1] to 310[4] in the input feature map data 310, this data will be read repeatedly during the convolution operation. Therefore, for the input feature map data 310, the total amount of data to be read can be derived from the following equation (4):
[0044] (w2 / tile_w2)×(h2 / tile_h2)×tile_h1×tile_w1×c1…(4)
[0045] In equations (2) to (4), the width f_w, height f_h, width stride stride_w, height stride stride_h, width w2, height h2, number of channels c1, and number of channels c2 are fixed values in the convolutional neural network model, and the capacity X of the memory 232 can be known in advance. Therefore, equations (1) to (3) can be used to find the width tile_w1 and height tile_h1 (corresponding to the data size of the input feature map data 310) and the width tile_w2 and height tile_h2 (corresponding to the data size of the output feature map data 320) that satisfy equation (1) and make equation (4) have the minimum value. It should be understood that when equation (1) is satisfied and equation (4) can have the lowest value, it means that the access times can be reduced as much as possible under the condition of meeting the capacity of the memory 232. In this way, the bandwidth requirements of the memory 220 and / or the memory 232 can be reduced.
[0046] Continue to refer to Figure 4 , in step S420, according to the second memory (for example Figure 2the capacity of the memory 233) and a first portion of data (e.g., but not limited to, block data of a row) of the last output feature map data among the multiple output feature map data (e.g., Figure 3B determine the number of levels of the multi-level operation and the amount of data required for at least a second portion of data (e.g., but not limited to, block data of a row) of each of the remaining data (e.g., Figure 3B the multiple output feature map data 330-1 to 330-(n-1)) among the multiple output feature map data. During the execution of the multi-level operation, the memory management circuit stores the first portion of data and the at least a second portion of data in the second memory.
[0047] To illustrate step S420, please refer to Figure 5A and Figure 5B . Figure 5A is a conceptual schematic diagram of step S420 in Figure 4 drawn according to some embodiments of the present application. In some embodiments, step S420 can be used to improve the hit rate of reading the memory 233 and reduce the number of accesses to the memory 220. As previously described, for each level of operation, the data size of the input feature map data is different from the data size of the output feature map data. To enable the memory 233 to store as many block data as possible, the number of levels of the multi-level operation can be determined by a backtracking method.
[0048] For example, as Figure 5AAs shown, the data size and required data volume (equivalent to the aforementioned specific preset value) of the block data of the input feature map data of the nth-level operation (equivalent to the output feature map data 330-(n-1) generated by the (n-1)th-level operation) can be estimated by using the aforementioned formulas (2) and (3) and based on a first part of the data (for example, but not limited to, the block data of a row) in the output feature map data 330-n of the last level (i.e., the nth level). Then, the data size and required data volume of the block data of the input feature map data of the (n-1)th-level operation (equivalent to the output feature map data 330-(n-2) generated by the (n-2)th-level operation) can be estimated again by using the aforementioned formulas (2) and (3) and based on at least a second part of the data (for example, the block data of a row) in the output feature map data 330-n of the (n-1)th level. And so on, until the data size and required data volume of the block data in the input feature map data of the first-level operation (equivalent to the aforementioned first preset value) are obtained. Then, the data volume of the first part of the data and the required data volume for generating at least a second part of the data in the remaining-level operations can be summed up to obtain a total data volume, and it can be confirmed whether the total data volume exceeds the capacity of the memory 233. If the total data volume does not exceed the memory 233, the number of levels (i.e., the value n) is incremented by 1 and the estimation is performed again. Or, if the total data volume exceeds the memory 233, the number of levels is set to n-1.
[0049] Figure 5B Drawn in accordance with some embodiments of the present application Figure 4Detailed flowchart of step S420 in []. In step S501, the capacity of a second memory (such as memory 233) is obtained. In step S502, the data sizes of the block data of the input feature map data and the output feature map data in each level of operation are calculated. For example, as previously described, the data sizes of the block data in the input feature map data and the output feature map data used in each level of operation can be calculated using the aforementioned equations (2) and (3). In step S503, it is assumed that the number of levels of the multi-level operation is a first value (such as value n), and the first value is greater than or equal to 2. In step S504, based on the first part of the data (such as, but not limited to, the block data of one row) in the output feature map data of the last level of operation, the data amount required to back-calculate at least one second part of the data for each of the remaining data in the multiple output feature map data is determined. In step S505, the data amount of the first part of the data and the data amount required to generate at least one second part of the data for each of the remaining data in the multiple output feature map data are summed to obtain a total data amount, and it is confirmed whether the total data amount is greater than the capacity of the second memory. If the total data amount is greater than the capacity of the second memory, it is confirmed that the number of levels is the first value minus 1, and if the total data amount is less than the capacity of the second memory, the number of levels is updated to a second value, and steps S504 and S505 are executed again, where the second value is the first value plus 1.
[0050] Through multiple S501 - S505, one layer of convolution operation of the convolutional neural network model can be split into multiple operations. Therefore, based on the same concept, by executing step S420 multiple times, multiple layers of convolution operations of the convolutional neural network model can be further split into multi-level operations.
[0051] Continue to refer to Figure 4 In step S430, record the data size of the block data, the number of levels of the multi-level operation, the data amount of the first part of the data, and the data amount of the at least one second part of the data as a predetermined compilation file, where the memory management circuit accesses the dynamic random access memory, the first memory, and the second memory based on the predetermined compilation file.
[0052] As previously described, through step S420, each layer of convolution operation (corresponding to different instructions) of the convolutional neural network model can be split into multiple operations respectively. In this way, the correspondence between various information obtained through steps S410 and S420 (such as the number of levels of the multi-level operation, the data sizes and required data amounts of the block data in the input feature map data and the output feature map data used in the multi-level operation, etc.) and multiple instructions can be recorded as a predetermined compilation file. In this way, Figure 2The processor 210 can issue an instruction CMD according to this predetermined compilation file, and the memory management circuit 231 can determine how to split the convolution operation corresponding to the instruction CMD based on the instruction CMD, and access the memory 220, the memory 232, and the memory 233 accordingly.
[0053] In some embodiments, the method 400 for reducing memory bandwidth can be executed by a computer-aided design system and / or a circuit simulation software to generate the predetermined compilation file, and the predetermined compilation file can be pre-stored in a buffer (not shown) of the artificial intelligence system 200. Thus, the processor 210 can issue an instruction CMD according to the predetermined compilation file. In other embodiments, the method 400 for reducing memory bandwidth can also be executed by the processor 210. The above application manners of the method 400 for reducing memory bandwidth are only examples, and the present application is not limited thereto.
[0054] Figure 4 and Figure 5B The multiple operations and / or steps of are only examples and do not limit that they need to be executed in the order in this example. Without departing from the operation manners and scopes of the embodiments of the present application, in Figure 4 and Figure 5B each operation and / or step of can be appropriately increased, replaced, omitted, or executed in a different order (for example, they can be executed simultaneously or partially simultaneously).
[0055] In summary, the intelligent processor device and the method for reducing memory bandwidth in some embodiments of the present application can split the multi-layer convolution operations of the convolutional neural network model into multi-stage operations, and temporarily store the operation results generated during the execution of the multi-stage operations in an additional cache memory. In this way, the access times and data access volume of the original memory of the system can be reduced, so as to reduce the bandwidth requirement of the memory.
[0056] Although the embodiments of the present application are as described above, these embodiments are not used to limit the present application. Those of ordinary skill in the art can make changes to the technical features of the present application based on the explicit or implicit content of the present application. All such changes may fall within the scope of patent protection sought by the present application. In other words, the scope of patent protection of the present application shall be subject to what is defined in the scope of patent application of this specification.
[0057] Symbol description:
[0058] 200: Artificial intelligence system;
[0059] 210: Processor;
[0060] 220, 232, 233: Memory;
[0061] 230: Intelligent processor device;
[0062] 231: Memory management circuit;
[0063] 234: Convolution operation circuit;
[0064] 310: Input feature map data;
[0065] 310[1], 310[2], 310[3], 310[4]: Block data;
[0066] 315: Convolution kernel;
[0067] 320, 320-1, 320-2, 320-3, 320-4, 330-1, 330-2, 330-3, 330-4, 330-n: Output feature map data;
[0068] 320[1], 320[2], 320[3], 320[4]: Block data;
[0069] 400: Method for reducing memory bandwidth;
[0070] CMD: Instruction;
[0071] DI: Input data;
[0072] DO: Output data;
[0073] S3-11, S3-12, S3-13, S3-14, S3-15, S3-16: Steps;
[0074] S3-21, S3-22, S3-23, S3-24, S3-25, S3-26: Steps;
[0075] S3-n1, S3-n2, S3-n3, S3-n4, S3-n5, S3-n6: Steps;
[0076] S410, S420, S430: Steps;
[0077] S501, S502, S503, S504, S505: Steps;
[0078] f_h, h1, h2: Height;
[0079] f_w, w1, w2: Width.
Claims
1. A method for reducing memory bandwidth, characterized in that, An intelligent processor device applied to execute a convolutional neural network model, wherein the intelligent processor device includes a first memory, a memory management circuit, a second memory, and a convolutional operation circuit, and the method for reducing memory bandwidth includes: Determine the data size of a block of data stored in the first memory when the convolutional operation circuit performs a convolutional operation according to the capacity of the first memory, wherein the memory management circuit transfers the block of data from a dynamic random access memory to the first memory, and the convolutional operation circuit sequentially performs multiple levels of operations corresponding to the convolutional operation on the block of data to sequentially generate multiple output feature map data; Determine the number of levels of the multiple levels of operations and the required data amount of at least a second part of data for each of the remaining data in the multiple output feature map data according to the capacity of the second memory and the data amount of a first part of data of the last output feature map data in the multiple output feature map data, wherein during the execution of the multiple levels of operations, the memory management circuit stores the first part of data and the at least a second part of data in the second memory; and Generate a predetermined compilation file based on the data size of the block of data, the number of levels of the multiple levels of operations, the data amount of the first part of data, and the data amount of the at least a second part of data, wherein the memory management circuit accesses the dynamic random access memory, the first memory, and the second memory based on the predetermined compilation file.
2. The method according to claim 1, wherein The data amount of the at least a second part of data is sufficient for the convolutional operation circuit to generate the data amount of the first part of data.
3. The method according to claim 1, characterized in that Determine the data size of the block of data stored in the first memory when the convolutional operation circuit performs the convolutional operation according to the capacity of the first memory includes: Determine the data size of the block of data according to the capacity of the first memory, the data size of an input feature map data of the convolutional neural network model, the data size of an output feature map data, and the data size of a convolutional kernel data.
4. The method according to claim 1, wherein Determine the number of levels of the multiple levels of operations and the data amount of the at least a second part of data for each of the remaining data in the multiple output feature map data according to the capacity of the second memory and the data amount of the first part of data of the last output feature map data in the multiple output feature map data includes the following steps: (a) Calculate the data size of a block of data in the multiple input feature map data used in each of the multiple levels of operations and the data size of a block of data in the multiple output feature map data; (b) Assume that the number of levels of the multiple levels of operation data is a first value, wherein the first value is a positive integer greater than or equal to 2; (c) Back-infer the amount of data required for at least one second part of each of the remaining data in generating the multiple output feature map data based on the first part of the data of the last output feature map data, wherein the last output feature map data corresponds to the last stage of the multi-stage operation; and (d) Sum the amount of data of the first part of the data and the amount of data required for at least one second part of each of the remaining data in generating the multiple output feature map data to obtain a total amount of data, and confirm whether the total amount of data is greater than the capacity of the second memory. If the total amount of data is greater than the capacity of the second memory, determine that the number of stages of the multi-stage operation is the first value, and if the total amount of data is less than the capacity of the second memory, update the number of stages of the multi-stage operation to a second value, and execute step (c) and step (d) again, wherein the remaining data in the multiple output feature map data corresponds to the remaining stages of the multi-stage operation, and the second value is the first value plus 1.
Citation Information
Patent Citations
Convolutional neural network-based image processing method and device, and unmanned aerial vehicle
US20210192246A1
Operation Accelerator, Processing Method, and Related Device
US20210224125A1