Artificial intelligence algorithm computing acceleration processor and method, system, and readable medium
By setting up multiple temporary storage areas and operators in the memory unit, convolution operations are performed in stages, and the second operator is triggered when the first calculation result reaches a predetermined data amount, the problem of low computing efficiency between single data reading and storage is solved, and efficient convolution operations and low power consumption are realized.
Patent Information
- Application Number
- CN202210008974.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-11-08
- Filing Date
- 2022-01-06
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-06
AI Technical Summary
The existing convolutional operation method is not efficient between single data reading and storage, resulting in large power consumption and poor operator utilization.
An artificial intelligence algorithm operation acceleration processor is proposed. By setting up a plurality of temporary storage areas and operators in the memory unit, the calculation is performed in stages, and the second operator is triggered to perform further operations when the first calculation result reaches a predetermined data amount.
The efficiency of convolutional operations is improved, the number of read and write times to memory cells is reduced, power consumption is reduced, and the utilization rate of the calculator is improved.
Smart Images

Figure CN114819117B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an artificial intelligence (AI) algorithm computing acceleration processor and method thereof, a computer system and a non-transitory computer-readable medium. Background Art
[0002] Edge computing is a network computing architecture that places the computing process as close to the data source as possible to reduce latency and bandwidth usage. The goal is to reduce the amount of computing performed in a centralized remote location (such as the "cloud"), thereby minimizing the amount of communication between remote clients and servers. In recent years, the rapid development of technology has made edge computing more practical.
[0003] In edge computing, user terminal devices (such as but not limited to smartphones) can speed up data processing and transmission and reduce latency. Edge computing can be accomplished by relying on AI hardware accelerators in user terminal devices.
[0004] Neural networks have made great breakthroughs in recent years, from the initial Perceptron to AlexNet and VGG, and their accuracy has been continuously improved, but the model has become more complex. Overly complex models make the amount of calculation so large that they cannot be run on low-end products (such as mobile phones). MobileNet was invented to solve this problem and improve processing speed.
[0005] The core idea of MobileNet is to simplify the traditional convolution process into a process with less computation, mainly by dividing the operation process into depthwise convolution and pointwise convolution.
[0006] As for MobileNet V1, its accuracy is good and it can greatly improve the processing speed. MobileNet V1 introduces deep convolution instead of traditional standard convolution to reduce the amount of calculation. In order to further improve it, MobileNet V2 has been proposed.
[0007] Compared with MobileNet V1, MobileNet V2 mainly introduces two changes: Linear Bottleneck and Inverted Residual Blocks.
[0008] The linear bottleneck removes the nonlinear enabling layer after the small-dimensional output layer in order to ensure the expressiveness of the model.
[0009] Inverted Residual Blocks: In residual blocks, the dimension is first reduced and then expanded, while in inverted residual blocks, the opposite is true, the dimension is first expanded and then reduced. The advantage is that the reused features can be reused to alleviate the degradation of features.
[0010] In order to improve the traditional convolution operation method, all parties are thinking about improving it to propose efficient convolution operations. However, the traditional convolution operation method is to read the input data from the memory unit, and the operator performs a single operation and then stores it back to the memory unit. The data is read, calculated and stored repeatedly according to the algorithm used. Each time the data is read and stored in the memory unit, there will be power consumption problems. Therefore, how to maximize the operation between a single data read and storage is a major issue in efficient convolution operations. At the same time, the focus of efficient convolution operations is to improve the traditional convolution operation method into several stages, and the amount of operations required in these stages is different, which in turn causes the problem of poor utilization rate of the same operator in different stages.
[0011] Therefore, how to develop an artificial intelligence (AI) algorithm computing acceleration processor and method thereof, a computer system and a non-transitory computer-readable medium that have both high performance and low power consumption is one of the directions of the industry's efforts. Summary of the invention
[0012] According to an embodiment of the present case, an artificial intelligence algorithm operation acceleration processor is proposed, which is suitable for operating an input data in a memory unit. The memory unit includes: a first data storage area for storing the input data; a second data storage area for storing a description data, wherein the description data includes a weight data; and a third data storage area for storing an output data. The artificial intelligence algorithm operation acceleration processor includes: a first temporary storage area for temporarily storing a first part of the input data, wherein the first temporary storage area has a predetermined data length; a second temporary storage area for temporarily storing a first part of the description data; a third temporary storage area for temporarily storing a first part of the weight data; a first operator for operating the first part of the input data and the first part of the weight data to generate a first operation result; a fourth temporary storage area for temporarily storing the first operation result; a fifth temporary storage area for temporarily storing a second part of the weight data; and a second operator for operating the first operation result and the second part of the weight data to generate a second operation result. When the fourth temporary storage area is full of a predetermined amount of data, the second operator is triggered to operate the first operation result and the second part of the weight data.
[0013] According to another embodiment of the present case, a method for accelerating computational processing of an artificial intelligence algorithm is proposed, comprising the following steps: A reading an input data and a description data from a memory unit, wherein the description data also includes a weight data; B using a first operator to operate a first part of the input data and a first part of the weight data to generate a first computation result; C temporarily storing the first computation result; D triggering a second operator when the first computation result stores a predetermined amount of data, and using the second operator to operate the first computation result and a second part of the weight data to generate a second computation result; and E writing the second computation result to the memory unit.
[0014] According to another embodiment of the present invention, a computer system is provided, comprising:
[0015] A memory unit includes: a first data storage area for storing input data; a second data storage area for storing description data, wherein the description data includes weight data; and a third data storage area for storing output data;
[0016] a memory read / write controller, coupled to the memory unit, to control reading and writing of the memory unit; and
[0017] An artificial intelligence algorithm computing acceleration processor is coupled to the memory read / write controller, and the artificial intelligence algorithm computing acceleration processor includes:
[0018] a first temporary storage area for temporarily storing a first part of the input data, wherein the first temporary storage area has a predetermined data length;
[0019] a second temporary storage area, for temporarily storing a first part of the description data;
[0020] a third temporary storage area, used for temporarily storing a first part of the weight data;
[0021] a first operator, for operating the first part of the input data and the first part of the weight data to generate a first operation result;
[0022] a fourth temporary storage area, used for temporarily storing the first operation result;
[0023] a fifth temporary storage area, for temporarily storing a second part of the weight data; and
[0024] a second operator, used for operating the first operation result and the second part of the weight data to generate a second operation result;
[0025] When the fourth temporary storage area is full of a predetermined amount of data, the second operator is triggered to operate the first operation result and the second part of the weight data.
[0026] According to another embodiment of the present invention, a non-transitory computer-readable medium is provided, wherein a program code that can be read and executed by a computer is stored. When the program code is executed, the computer performs the following actions:
[0027] A reads an input data and a description data from a memory unit, wherein the description data further includes a weight data;
[0028] B. operating a first part of the input data and a first part of the weight data with a first operator to generate a first operation result;
[0029] C temporarily stores the first operation result;
[0030] D. When the first operation result stores a predetermined amount of data, triggering a second operation unit to operate the first operation result and a second part of the weight data by the second operation unit to generate a second operation result; and
[0031] E writes the second operation result into the memory unit.
[0032] In order to better understand the above and other aspects of the present invention, embodiments are given below and described in detail with reference to the accompanying drawings: BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 A functional block diagram of a computer system according to an embodiment of the present invention is shown.
[0034] Figure 2 A flow chart of an artificial intelligence algorithm accelerated computing method according to an embodiment of the present invention is shown.
[0035] Figure 3A and Figure 3B A flow chart of an artificial intelligence algorithm accelerated computing method according to another embodiment of the present invention is shown.
[0036] Figure 4A A schematic diagram showing the operation of a first operating unit according to an embodiment of the present invention is shown.
[0037] Figure 4B A schematic diagram showing the operation of the second operator according to an embodiment of the present invention is shown.
[0038] Figure 5 The flowchart of writing data into the fourth temporary storage memory according to one embodiment of the present invention is shown.
[0039] Figure 6The input data in the input data storage area of the memory unit according to one embodiment of the present invention is displayed.
[0040] Fig. 7A Displaying the first part of the weight data according to an embodiment of the present case, Figure 7B A second portion of weight data according to an embodiment of the present invention is shown.
[0041] FIG. 8A to FIG. 8H A schematic diagram showing the operation of an artificial intelligence algorithm computing acceleration processor according to an embodiment of the present invention.
[0042] Fig. 9 In one embodiment of the present invention, the data output is shown when the movement amount of the first layer convolution operation is 1 and 2 respectively.
[0043] Description of Reference Numerals
[0044] 100: Computer Systems
[0045] 110: memory unit
[0046] 111: Input data storage area
[0047] 112: Describe the data storage area
[0048] 113: Output data storage area
[0049] 115: Memory read / write controller
[0050] 120: Artificial Intelligence Algorithm Computing Acceleration Processor
[0051] 121: First temporary storage area
[0052] 122: Second temporary storage area
[0053] 123: The third temporary storage area
[0054] 124: First operator
[0055] 125: Fourth temporary storage area
[0056] 126: Fifth temporary storage area
[0057] 127: Second operator 127
[0058] 128:Activation unit
[0059] 129: Pooling unit
[0060] 210-280: Steps
[0061] 302-340: Steps
[0062] 401: Input Data
[0063] 402: The first part of the weight data
[0064] 411, 412: Data
[0065] 421, 422: Second operation result
[0066] 510: Data DETAILED DESCRIPTION
[0067] The technical terms in this specification refer to the customary terms in the technical field. If some terms are explained or defined in this specification, the interpretation of these terms shall be based on the explanation or definition in this specification. Each embodiment of the present disclosure has one or more technical features. Under the premise of possible implementation, a person with ordinary knowledge in the technical field can selectively implement some or all of the technical features in any embodiment, or selectively combine some or all of the technical features in these embodiments.
[0068] Please refer to Figure 1 , which illustrates a functional block diagram of a computer system 100 according to an embodiment of the present invention. The computer system 100 includes a memory unit 110, a memory read / write controller 115, and an artificial intelligence algorithm operation acceleration processor 120. The memory read / write controller 115 is coupled to the memory unit 110 and the artificial intelligence algorithm operation acceleration processor 120.
[0069] The artificial intelligence algorithm operation acceleration processor 120 is adapted to operate on an input data in the memory unit 110 (eg, a dynamic random access memory DRAM).
[0070] The memory unit 110 includes: an input data storage area 111 for storing input data IN; a description data storage area 112 for storing description data (descriptor), wherein the description data includes weight data; and an output data storage area 113 for storing output data.
[0071] The memory read / write controller 115 reads data (such as input data IN and description data) from the memory unit 110 to the artificial intelligence algorithm operation acceleration processor 120 so that the artificial intelligence algorithm operation acceleration processor 120 performs a MAC (Multiply Accumulate, MAC) operation, and the memory read / write controller 115 writes the MAC operation result of the artificial intelligence algorithm operation acceleration processor 120 to the memory unit 110.
[0072] The artificial intelligence algorithm operation acceleration processor 120 includes: a first temporary storage area 121 (for example, a static random access memory SRAM) for temporarily storing part of the input data, wherein the first temporary storage area 121 has a predetermined data length; a second temporary storage area 122 (for example, a static random access memory SRAM) for temporarily storing part of the description data; a third temporary storage area 123 (for example, a static random access memory SRAM) for temporarily storing a first part of the weight data; a first operator 124 (for example, a multiplication and accumulation operator) for calculating the input data and the first part of the weight data to generate a first operation result, wherein the The first operator has a first maximum operation amount; a fourth temporary storage area 125 (for example, a static random access memory SRAM) for temporarily storing the first operation result, wherein the fourth temporary storage area includes at least three times (or more than three times) the predetermined data length; a fifth temporary storage area 126 (for example, a static random access memory SRAM) for temporarily storing a second part of the weight data; a second operator 127 (for example, a multiplication and accumulation operator) for calculating the first operation result and the second part of the weight data to generate a second operation result, wherein the second operator has a second maximum operation amount, and the second maximum operation amount is less than the first maximum operation amount. When the fourth temporary storage area 125 is full of a predetermined data amount, the second operator 127 is triggered to calculate the first operation result and the second part of the weight data. When the second operator 127 is calculating, the first operator 124 continues to calculate the input data. The setting of the predetermined data amount is determined according to the description data. Furthermore, the setting of the predetermined data amount for triggering the second operator is determined according to a batch width and a feature extractor parameter.
[0073] In a possible embodiment of the present case, the artificial intelligence algorithm operation acceleration processor 120 further selectively includes an activation unit 128 (Activation) for performing an activation operation on the first operation result output by the first operator 124. The operations that the activation unit 128 can perform include, but are not limited to, rectified linear unit function (Rectified Linear Unit, ReLU), sigmoid function (Sigmoid), hyperbolic tangent function (Tanh), etc. In an embodiment of the present case, the activation operation is a selective operation, which is set in the description data.
[0074] In a possible embodiment of the present case, the artificial intelligence algorithm operation acceleration processor 120 further selectively includes a pooling unit 129 (Pooling) for performing a pooling operation on the first operation result output by the fourth temporary storage area 125. The operations that the pooling unit 129 can perform include, but are not limited to: Max-Pooling, Mean-Pooling, Stochastic-Pooling, etc. The pooling operation result of the pooling unit 129 is input to the memory read-write controller 115. In an embodiment of the present case, the pooling operation and the second operation are at the same level, and one can be selected and set in the description data.
[0075] In a possible embodiment of the present case, the first operator 124 further includes a first operation unit array, the first operation unit array includes a plurality of first operation units, each of the first operation units (element) is configured to: receive the first part of the input data and the weight data corresponding to a multi-dimensional position; and process the first part of the input data and the weight data, and generate a plurality of operation result values as the first operation result. In the present embodiment, "multi-dimensional position" refers to different data points, such as but not limited to data at each coordinate point in a two-dimensional plane coordinate system.
[0076] In a possible embodiment of the present case, the second operator 127 further includes a second operation unit array, which includes a plurality of second operation units, each of which is configured to: receive the first operation result and the second part of the weight data corresponding to the multi-dimensional position; and process the first operation result and the second part of the weight data and generate a plurality of operation result values as the second operation result. The second operation result generated by the second operator 127 can be written to the memory unit 110 through the memory read-write controller 115. The number of the first operation units is greater than the number of the second operation units. When the second operator is operating, the first operator and the second operator are in a parallel processing state, and the parallel processing state means that the first operator and the second operator perform their respective operation processing work at the same time.
[0077] In one embodiment of the present case, the description data includes, but is not limited to: layer number, feature extractor settings, pooling settings, input feature map size, number of channels, input feature map starting address, output feature map starting address, sub-layer description file pointer, activation settings, etc.
[0078] In an embodiment of the present case, the first temporary storage area 121 is, for example but not limited to, a first-in first-out (FIFO) temporary storage area, and sends input data to the first operator 124 in a first-in first-out manner.
[0079] Please refer to Figure 2 , which illustrates an artificial intelligence algorithm accelerated computing processing method according to an embodiment of the present case, comprising: reading a first portion of an input data from a memory unit to a first temporary storage area (210); reading a description data from the memory unit to a second temporary storage area, wherein the description data includes a weight data (220); reading a first portion of the weight data from the memory unit to a third temporary storage area (230); reading a second portion of the weight data from the memory unit to a fifth temporary storage area (240); reading the input data from the first temporary storage area and reading the first portion of the weight data from the third temporary storage area, with a first computing The device performs a first operation and obtains a first operation result (250); writes the first operation result to a fourth temporary storage area (260); when the first operation result in the fourth temporary storage area reaches a predetermined data amount, (1) reads the first operation result from the fourth temporary storage area and reads out the second part of the weight data from the fifth temporary storage area, performs a second operation with a second operator and obtains a second operation result, or, (2) performs a pooling operation on the first operation result output by the four temporary storage memories to obtain a pooling operation result (270); and writes the second operation result or the pooling operation result to the memory unit (280).
[0080] In one embodiment of the present invention, a selective activation operation may be included between steps 250 and 260 .
[0081] Figure 3A FIG3B is a flow chart of an artificial intelligence algorithm accelerated computing method according to another embodiment of the present invention.
[0082] In step 302, the artificial intelligence algorithm operation acceleration processor 120 reads the description data from the description data storage area 112 of the memory unit 110. In detail, after the system writes the input data and the description data into the memory unit 110, the system will send a notification to the artificial intelligence algorithm operation acceleration processor 120 to allow the artificial intelligence algorithm operation acceleration processor 120 to read the input data and the description data. In this way, the artificial intelligence algorithm operation acceleration processor 120 can be triggered to perform operations.
[0083] In step 304, the artificial intelligence algorithm operation acceleration processor 120 reads a section of input data from the input data storage area 111 of the memory unit 110 to the first temporary storage area 121. This section starts from the memory address I (h, w) (h and w are both positive integers), and the width of the read data is the section width sect_width.
[0084] In step 306 , the artificial intelligence algorithm operation acceleration processor 120 reads the first part of the weight data from the description data storage area 112 of the memory unit 110 to the third temporary storage area 123 .
[0085] In step 307, it is determined whether h≧(ft_size 1st -1), and (h% Stride 1st == 0)", where h%Stride 1st == 0" indicates whether the data address h is 1st Integer divisibility, ft_size 1st is the feature extractor size of the first convolution operation, Stride 1st Represents the movement amount of the first layer of convolution operation. In the convolution operation, the feature extractor (filter) (or kernel) is used as the basis, and the operation target is moved one by one, and the parameter "Stride" is set for the movement amount of the feature extractor. When the parameter "Stride" is set to 1, it means that an operation is performed every time it moves forward; when the parameter "Stride" is set to 2, it means that an operation is performed only once every two times it moves forward. Therefore, when the parameter "Stride" is set to 2 or more, the amount of operation can be reduced. In the embodiment of the present case, this is an optional approach. If step 307 is yes, the process continues to step 308; and if step 307 is no, the process continues to step 318. For example, when the feature extractor size of the first layer of convolution operation is 1, only the input data of h=0 needs to be read in. When the feature extractor size of the first layer of convolution operation is 3, it is necessary to read the data of h=0, h=1 and h=2 in the input data of the block before step 308 is performed.
[0086] In step 308 , the AI algorithm operation acceleration processor 120 loads a batch of input data from the first temporary storage area 121 to the first operator 124 , wherein the data width of the batch is a batch width WB (WB is a positive integer), and the batch width is less than or equal to the block width.
[0087] In step 310 , the first operator 124 of the artificial intelligence algorithm operation acceleration processor 120 operates the first part of the input data and the weight data to generate a first operation result.
[0088] In step 312, the first operator 124 of the artificial intelligence algorithm operation acceleration processor 120 writes the first operation result to the fourth temporary storage area 125. For example but not limited to, the fourth temporary storage area 125 has at least m times the predetermined data length (for example but not limited to, m=3), and the fourth temporary storage area 125 is overwritable, wherein the predetermined data length is, for example, equal to the block width.
[0089] In step 314, it is determined whether all the input data of the block in the first temporary storage area 121 have been read out and calculated. If step 314 is no, the process returns to step 308, and the artificial intelligence algorithm calculation acceleration processor 120 loads the next batch of input data (data width WB) from the first temporary storage area 121 to the first operator 124. If step 314 is yes, the process continues to step 316.
[0090] In step 316, it is determined whether all data in the fourth temporary storage area 125 have been processed, for example but not limited to, determining whether h is equal to h max , where h max is the maximum value of the data address h of the input data. If step 316 is no, the flow proceeds to step 318; and, if step 316 is yes, the flow proceeds to step 320.
[0091] In step 318, the parameter h is updated, for example, the parameter h is updated to h=h+1, so as to read the data of the next row.
[0092] In step 320, it is determined whether there is any input data in the first temporary storage area 121 that has not been read out. If the answer to step 320 is no (i.e., all the input data in the first temporary storage area 121 has been read out), the operation flow ends. If the answer to step 320 is yes (i.e., there is still input data in the first temporary storage area 121 that has not been read out), the flow continues to step 322.
[0093] In step 322, the parameter w is updated and the parameter h is reset, for example but not limited to, the parameter w is updated to w=w+sect_width-(ft_size 1st -1 +ft_size 2nd -1), and reset the parameter h to h=0, where ft_size 2ndis the feature extractor size for the second convolution operation. After step 322 is executed, the process returns to step 304. In one embodiment of the present case, for example, if sect_width is 32, a block (section) of input data will be read from the input data storage area 111 in the memory unit 110 for the first time, and the data from the 1st (address 0) to the 32nd (address 31) of the input data will be read; and during the second operation, the starting address of the read data will be determined according to the feature extractor size used in the operation. The setting of the feature extractor size is stored in the description data, for example but not limited to, the first layer feature extractor (ft_size 1st ) is 1*1, the second layer feature extractor (ft_size 2nd ) is 3*3. Since the first data operation of the second layer at this time needs to use the data of the 31st (address 30) to the 33rd (address 32) to calculate, it is necessary to read a block (section) of input data from the input data storage area 111 in the memory unit 110, and read the data of the 31st (address 30) to the 62nd (address 61) of the input data for calculation.
[0094] In addition, after executing step 312, step 324 is further executed.
[0095] In step 324, it is determined whether the stored data (ie, the first operation result) written into the fourth temporary storage area 125 meets a predetermined data amount. If step 324 is yes, the process continues to step 326; and if step 324 is no, the process continues to step 335.
[0096] In step 326, it is determined whether 1st % Stride 2nd == 0". If step 326 is yes, the process continues to step 328; and, if step 326 is no, the process continues to step 335. Step 326 is an optional step like step 307. 1st % Stride 2nd == 0" means h 1st Is it divisible by the parameter stride2nd, where Stride 2nd Represents the movement of the second layer convolution operation, h 1st It is the data address h value of the first operation result stored in the fourth temporary storage area 125.
[0097] In step 328, the data of the fourth temporary storage area 125 is read to the second operator 127 according to the size of the second-layer feature extractor. For example, if the size of the second-layer feature extractor is 3*3, the data of the address "p([0..2], [w..w+2])" of the fourth temporary storage area 125 is read to the second operator 127. In another embodiment, if the size of the second-layer feature extractor is 5*5, the data of the address "p([0..4], [w..w+4])" of the fourth temporary storage area 125 is read to the second operator 127.
[0098] In addition, in step 330, the AI algorithm operation acceleration processor 120 reads the second part of the weight data from the description data storage area 112 of the memory unit 110 to the fifth temporary storage area 126. In one embodiment of the present case, step 330 can be completed simultaneously with step 304 and step 306.
[0099] In step 332, the second operator 127 of the artificial intelligence algorithm operation acceleration processor 120 calculates the first operation result (i.e., the data read from the fourth temporary storage area 125 in step 328) and the second part of the weight data (i.e., the data stored in the fifth temporary storage area 126 in step 330) to generate a second operation result.
[0100] In step 334 , the second operation result of the second operator 127 is written into the memory unit 110 via the memory read / write controller 115 .
[0101] In step 335 , it is determined whether the data currently being calculated is the data of the first batch, for example, whether the w value is less than or equal to the batch width value. If step 335 is yes, the process continues to step 340 ; if step 335 is no, the process ends.
[0102] In step 336, it is determined whether all the data in the fourth temporary storage area 125 have been calculated by the second operator 127, for example, it is determined whether the value of w is equal to w max , where w max is the maximum value of the data address w of the first operation result data. In one embodiment, w max If step 336 is yes, the process continues to step 340 ; if step 336 is no, the process continues to step 338 .
[0103] In step 338, update the parameter w (w = w + Stride 2nd ), and the process returns to step 328.
[0104] In step 340, update the parameter h 1st (h 1st =h 1st+1). After that, the process ends.
[0105] Please refer to Figure 4A , which shows a schematic diagram of a first computing unit according to an embodiment of the present invention. Figure 4A In the example, the parameter ochb represents the number of output channel batches, and the parameter k represents the number of channels of the input data. In one embodiment, the first layer operation adopts a point-by-point convolution operation architecture, and uses a 1*1 feature extractor to operate on the input data to convert the number of channels. The amount of operation of the first layer operation is 1*1*k*ochb. Figure 4A As shown, individual input data (one input data is shown as a box 401) and the first part of individual weight data (the first part of one weight data is shown as a box 402) are multiplied and then added to obtain a first operation result. The first operation result of each round is written into the fourth temporary storage area 125.
[0106] Figure 4B A schematic diagram showing the operation of the second operator according to an embodiment of the present invention is shown. Figure 4B As shown, when 3*3 (=9) records of data 411 are written into the fourth temporary storage area 125, the second operator 127 operates on the second layer of input data (i.e., the 3*3 (=9) records of data 411 stored in the fourth temporary storage area 125) and the second part of the weight data (stored in the fifth temporary storage area 126) to obtain a second operation result 421. The second operation result 421 is written into the output data storage area 113 of the memory unit 110. Thereafter, when 3*3 (=9) records of data 412 are written into the fourth temporary storage area 125, the second operator 127 operates on the second layer of input data (i.e., the 3*3 (=9) records of data 412 in the fourth temporary storage area 125) and another second part of the weight data to obtain a second operation result 422. The second operation result 422 is written into the output data storage area 113 of the memory unit 110.
[0107] Figure 5A schematic diagram showing a process of writing data to the fourth temporary storage area 125 according to an embodiment of the present invention. In the first round, the first operator 124 writes the first operation result with a bit number of "WB (for example but not limited to, 8 bits)" to the 0th data line of the fourth temporary storage area 125. Then, in subsequent rounds, the first operator 124 writes subsequent first operation results to the 0th data line of the fourth temporary storage area 125. When the 0th data line is full, the first operation result is written to the 1st data line; and, when the 1st data line is full, the first operation result is written to the 2nd data line. The length of each data line is, for example, equal to the block width of the input data of the read memory unit 110.
[0108] In addition, a predetermined data amount is determined according to the size of the second-layer feature extractor. For example, when the size of the second-layer feature extractor is 3*3, the predetermined data amount is when the 0th data line in the fourth temporary storage area 125 is full, the 1st data line is full, and the 2nd data line is full of 3 records. Figure 5 As shown, the 3*3 data 510 have been written into the fourth temporary storage area 125, and the second operation can be started, that is, the second operator 127 operates on the input data of the second layer (that is, the 3*3 (=9) data 510 of the fourth temporary storage area 125) and the second part of the weight data to obtain a second operation result.
[0109] Figure 6 The input data stored in the input data storage area 111 of the memory unit 110 according to one embodiment of the present invention is shown. In one example, for example but not limited to, the input data may be an input feature map, whose size is h*w*k, and in one embodiment, for example, 4*32*48, and when the data is stored as an address, the input data is represented by I(0,0,0)~I(3,31,47).
[0110] Fig. 7A The first part of the weight data according to an embodiment of the present invention is shown. In one embodiment, under a feature extractor of size 1*1, the amount of weight data is 1*1*k*n, where k is the number of channels of input data and n is the number of channels of output data. Fig. 7A In the embodiment k=48, n=16, Figure 7B The second part of the weight data according to one embodiment of the present invention is shown. In one embodiment, the weight data amount is 3*3*n under a feature extractor of size 3*3. Figure 7B In the embodiment n=16. Fig. 7A Medium, F 0 (0,0,0)~ F 15 (0,0,47) represents the first part of the weight data of an embodiment of the present invention. Figure 7BIn, f 0 (0,0)~ f 15 (2,2) shows the second part of the weight data of one embodiment of the present case.
[0111] FIG. 8A to FIG. 8H A schematic diagram showing the operation of the artificial intelligence algorithm acceleration processor 120 according to an embodiment of the present invention. Fig. 8A The first operation of the first round is shown. The operation descriptions of the third to sixth rounds are the same as those of the first and second rounds. In the following embodiment, the input data size is 4*32*48, the first layer feature extractor size is 1*1, the output channel is 16, the second layer feature extractor size is 3*3, the block width of the input data read each time is 32 (WS), and the batch data width of each round of operation of the first operator is 16 (WB).
[0112] A(0,n)=I(0,0,0)*F n (0,0,0)+I(0,0,1)*F n (0,0,1)+…+ I(0,0,47)*F n (0,0,47).
[0113] A(1,n)=I(0,1,0)*F n (0,0,0)+I(0,1,1)*F n (0,0,1)+…+I(0,1,47)*F n (0,0,47).
[0114] A(15,n)=I(0,15,0)*F n (0,0,0)+I(0,15,1)*F n (0,0,1)+…+ I(0,15,47)*F n (0,0,47).
[0115] P(0,0 . . . 15,n) includes: A(0,n)-A(15,n), wherein P(0,0 . . . 15,n) represents the first operation result of the first round written into the fourth temporary storage area 125 .
[0116] Figure 8B Displays the first operation of the second round.
[0117] A(0,n)=I(0,16,0)*F n (0,0,0)+I(0,16,1)*F n (0,0,1)+…+ I(0,16,47)*F n (0,0,47).
[0118] A(1,n)=I(0,17,0)*F n (0,0,0)+I(0,17,1)*F n (0,0,1)+…+ I(0,17,47)*F n (0,0,47).
[0119] A(15,n)=I(0,31,0)*F n (0,0,0)+I(0,31,1)*F n (0,0,1)+…+ I(0,31,47)*F n (0,0,47).
[0120] P(0,16 . . . 31, n) includes: A(0,n)-A(15,n), wherein P(0,16 . . . 31, n) represents the first operation result of the second round written into the fourth temporary storage area 125 .
[0121] Figure 8C Displays the first operation of the third round.
[0122] A(0,n)=I(1,0,0)*F n (0,0,0)+I(1,0,1)*F n (0,0,1)+…+ I(1,0,47)*F n (0,0,47).
[0123] A(1,n)=I(1,1,0)*F n (0,0,0)+I(1,1,1)*F n (0,0,1)+…+ I(1,1,47)*F n (0,0,47).
[0124] A(15,n)=I(1,15,0)*F n (0,0,0)+I(1,15,1)*F n (0,0,1)+…+ I(1,15,47)*F n (0,0,47).
[0125] P(1,0 . . . 15, n) includes: A(0, n) ˜A(15, n), wherein P(1,0 . . . 15, n) represents the first operation result of the third round written into the fourth temporary storage area 125 .
[0126] Fig.8D Displays the first operation of the fourth round.
[0127] A(0,n)=I(1,16,0)*F n(0,0,0)+I(1,16,1)*F n (0,0,1)+…+ I(1,16,47)*F n (0,0,47).
[0128] A(1,n)=I(1,17,0)*F n (0,0,0)+I(1,17,1)*F n (0,0,1)+…+ I(1,17,47)*F n (0,0,47).
[0129] A(15,n)=I(1,31,0)*F n (0,0,0)+I(1,31,1)*F n (0,0,1)+…+ I(1,31,47)*F n (0,0,47).
[0130] P(1,16 . . . 31, n) includes: A(0, n)-A(15, n), wherein P(1,16 . . . 31, n) represents the first operation result of the fourth round written into the fourth temporary storage area 125 .
[0131] Fig. 8E Displays the first operation of the fifth round.
[0132] A(0,n)=I(2,0,0)*F n (0,0,0)+I(2,0,1)*F n (0,0,1)+…+ I(2,0,47)*F n (0,0,47).
[0133] A(1,n)=I(2,1,0)*F n (0,0,0)+I(2,1,1)*F n (0,0,1)+…+ I(2,1,47)*F n (0,0,47).
[0134] A(15,n)=I(2,15,0)*F n (0,0,0)+I(2,15,1)*F n (0,0,1)+…+ I(2,15,47)*F n (0,0,47).
[0135] P(2,0 . . . 15, n) includes: A(0, n) ˜A(15, n), wherein P(2,0 . . . 15, n) represents the first operation result of the fifth round written into the fourth temporary storage area 125 .
[0136] Figure 8F-1 Displays the first operation in the sixth round, while Figure 8F-2 displays the second operation in the sixth round. The following explains Figure 8F-1 .
[0137] A(0,n) = I(2,16,0)*F n (0,0,0) + I(2,16,1)*F n (0,0,1) + … + I(2,16,47)*F n (0,0,47).
[0138] A(1,n) = I(2,17,0)*F n (0,0,0) + I(2,17,1)*F n (0,0,1) + … + I(2,17,47)*F n (0,0,47).
[0139] A(15,n) = I(2,31,0)*F n (0,0,0) + I(2,31,1)*F n (0,0,1) + … + I(2,31,47)*F n (0,0,47).
[0140] P(2,16…31,n) includes: A(0,n)~A(15,n), where P(2,16…31,n) represents the first operation result in the sixth round written to the fourth temporary storage area 125.
[0141] In the sixth round, since the size of the second-layer feature extractor is 3*3 and the first operation result in the fourth temporary storage area 125 has reached a predetermined data volume, the second operation can be started. That is, in an embodiment of the present case, when the quantity of the first operation result is sufficient, the second operation can be started. Instead of waiting until all the first operations are completed, written back to the memory unit, and then reading out the first operation result from the memory unit to start the second operation as in the prior art, it saves the time and power consumption of reading and writing in the memory unit. Especially in convolution operations, usually a large number of operation times are required, so the operation efficiency and power-saving function can be effectively improved.
[0142] Now explain Figure 8F-2 .
[0143] a(0,n) = P(0,0,n)*f n (0,0) + P(0,1,n)*f n (0,0) + P(0,2,n)*f n(0,0).
[0144] a(1,n)=P(1,0,n)*f n (1,0)+ P(1,1,n)*f n (1,1)+ P(1,2,n)*f n (1,2).
[0145] a(2,n)=P(2,0,n)*f n (2,0)+ P(2,1,n)*f n (2,1)+ P(2,2,n)*f n (2,2).
[0146] O(0,0,n)=a(0,n)+a(1,n)+a(2,n) O(0,0,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0147] Similarly,
[0148] a(0,n)=P(0,1,n)*f n (0,0)+ P(0,2,n)*f n (0,0)+ P(0,3,n)*f n (0,0).
[0149] a(1,n)=P(1,1,n)*f n (1,0)+ P(1,2,n)*f n (1,1)+ P(1,3,n)*f n (1,2).
[0150] a(2,n)=P(2,1,n)*f n (2,0)+ P(2,2,n)*f n (2,1)+ P(2,3,n)*f n (2,2).
[0151] O(0,1,n)=a(0,n)+a(1,n)+a(2,n) O(0,1,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0152] And so on.
[0153] a(0,n)=P(0,13,n)*f n (0,0)+ P(0,14,n)*f n (0,0)+ P(0,15,n)*f n (0,0).
[0154] a(1,n)=P(1,13,n)*f n (1,0)+ P(1,14,n)*f n (1,1)+ P(1,15,n)*f n (1,2).
[0155] a(2,n)=P(2,13,n)*f n (2,0)+ P(2,14,n)*f n (2,1)+ P(2,15,n)*f n (2,2).
[0156] O(0,13,n)=a(0,n)+a(1,n)+a(2,n) O(0,13,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0157] Figure 8G-1 The first operation of the seventh round is displayed, and Figure 8G-2 It shows that while the first operation of the seventh round is being performed, the second operation is being performed simultaneously. Figure 8G-2 The second operation is shown. Figure 8G-2 When the second operation is triggered, the second operation is independently executed, and the first operation is also executed simultaneously, and data is continuously stored in the fourth temporary storage area 125 for reading by the second operation.
[0158] A(0,n)=I(3,0,0)*F n (0,0,0)+I(3,0,1)*F n (0,0,1)+…+ I(3,0,47)*F n (0,0,47).
[0159] A(1,n)=I(3,1,0)*F n (0,0,0)+I(3,1,1)*F n (0,0,1)+…+ I(3,1,47)*F n (0,0,47).
[0160] A(15,n)=I(3,15,0)*F n (0,0,0)+I(3,15,1)*F n (0,0,1)+…+ I(3,15,47)*F n (0,0,47).
[0161] P(0,0 . . . 15, n) includes: A(0, n) ˜A(15, n), wherein P(0,0 . . . 15, n) represents the first operation result of the seventh round written into the fourth temporary storage area 125 .
[0162] Now explain Figure 8G-2 .
[0163] a(0,n)= P(0,14,n)*f n (0,0)+ P(0,15,n)*f n (0,0)+ P(0,16,n)*f n (0,0).
[0164] a(1,n)= P(1,14,n)*f n (1,0)+ P(1,15,n)*f n (1,1)+ P(1,16,n)*f n (1,2).
[0165] a(2,n)= P(2,14,n)*f n (2,0)+ P(2,15,n)*f n (2,1)+ P(2,16,n)*f n (2,2).
[0166] O(0,14,n)=a(0,n)+a(1,n)+a(2,n) O(0,14,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0167] Similarly,
[0168] a(0,n)= P(0,15,n)*f n (0,0)+ P(0,16,n)*f n (0,0)+ P(0,17,n)*f n (0,0).
[0169] a(1,n)= P(1,15,n)*f n (1,0)+ P(1,16,n)*f n (1,1)+ P(1,17,n)*f n (1,2).
[0170] a(2,n)= P(2,15,n)*f n (2,0)+ P(2,16,n)*f n (2,1)+ P(2,17,n)*f n (2,2).
[0171] O(0,15,n)=a(0,n)+a(1,n)+a(2,n) O(0,15,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0172] And so on.
[0173] a(0,n)= P(0,29,n)*f n (0,0)+ P(0,30,n)*f n (0,0)+ P(0,31,n)*f n (0,0).
[0174] a(1,n)= P(1,29,n)*f n (1,0)+ P(1,30,n)*f n (1,1)+ P(1,31,n)*f n (1,2).
[0175] a(2,n)= P(2,29,n)*f n (2,0)+ P(2,30,n)*f n (2,1)+ P(2,31,n)*f n (2,2).
[0176] O(0,29,n)=a(0,n)+a(1,n)+a(2,n) O(0,29,n) represents the (intermediate / final) output result to be written into the output data storage area 113 .
[0177] Figure 8H Displays the second operation as ongoing.
[0178] a(0,n)= P(0,14,n)*f n (0,0)+ P(0,15,n)*f n (0,0)+ P(0,16,n)*f n (0,0).
[0179] a(1,n)= P(1,14,n)*f n (1,0)+ P(1,15,n)*f n (1,1)+ P(1,16,n)*f n (1,2).
[0180] a(2,n)= P(2,14,n)*f n (2,0)+ P(2,15,n)*f n(2,1)+ P(2,16,n)*f n (2,2).
[0181] O(1,14,n)=a(0,n)+a(1,n)+a(2,n). O(1,14,n) represents the (intermediate / final) output result to be written to the output data storage area 113.
[0182] Similarly,
[0183] a(0,n)= P(0,15,n)*f n (0,0)+ P(0,16,n)*f n (0,0)+ P(0,17,n)*f n (0,0).
[0184] a(1,n)= P(1,15,n)*f n (1,0)+ P(1,16,n)*f n (1,1)+ P(1,17,n)*f n (1,2).
[0185] a(2,n)= P(2,15,n)*f n (2,0)+ P(2,16,n)*f n (2,1)+ P(2,17,n)*f n (2,2).
[0186] O(1,15,n)=a(0,n)+a(1,n)+a(2,n). O(1,15,n) represents the (intermediate / final) output result to be written to the output data storage area 113.
[0187] And so on.
[0188] a(0,n)= P(0,29,n)*f n (0,0)+ P(0,30,n)*f n (0,0)+ P(0,31,n)*f n (0,0).
[0189] a(1,n)= P(1,29,n)*f n (1,0)+ P(1,30,n)*f n (1,1)+ P(1,31,n)*f n (1,2).
[0190] a(2,n)= P(2,29,n)*f n (2,0)+ P(2,30,n)*fn (2,1)+ P(2,31,n)*f n (2,2).
[0191] O(1,29,n)=a(0,n)+a(1,n)+a(2,n). O(1,29,n) represents the (intermediate / final) output result to be written to the output data storage area 113.
[0192] Although the above example only describes up to the seventh round, those who are familiar with this technique should know how to extrapolate to subsequent rounds, and the details will not be repeated here.
[0193] In the aforementioned embodiment, when the second-layer feature extractor is 3*3, when wb=1 / 2ws, the second operation is triggered after the first operation of the fifth round is completed. In another embodiment, when the second-layer feature extractor is 5*5, when wb=(1 / 2)*ws, the second operation is triggered after the first operation of the ninth round is completed. In addition, in yet another embodiment, when the second-layer feature extractor is 3*3, when wb=(1 / 4)*ws, the second operation is triggered after the first operation of the ninth round is completed. In addition, in yet another embodiment, when the second-layer feature extractor is 3*3, when wb=1*ws, the second operation is triggered after the first operation of the third round is completed.
[0194] Fig. 9 According to one embodiment of the present invention, when the movement amount Stride of the first layer convolution operation is 1st When 1 and 2 are respectively, the data output is obtained. IFM stands for input feature map. Fig. 9 As shown, when the feature extractor size is 3*3 and the movement Stride 1st When it is 1 (the address of the next read block needs to be pushed forward 1 bit), the input data with a block width of 32 (WS) can be output after the second operation, O(0, 1, n)~O(0, 29, n), a total of 30 data outputs.
[0195] From the above description, it can be seen that in one embodiment of the present case, after performing several rounds, the first operation and the second operation can be performed simultaneously. Therefore, the embodiment of the present case has the advantage of better overall operation efficiency.
[0196] The above-mentioned embodiments of the present case are applicable to the algorithm architecture of efficient convolution operation to improve the disadvantage of the operator utilization rate of the existing convolution operation. As mentioned above, since the embodiments of the present case can integrate the phased operation of the efficient convolution operation into the approximately parallel processing, the operation efficiency can be effectively improved.
[0197] In addition, since the artificial intelligence algorithm computing acceleration processor of the embodiment of the present case can simultaneously have the advantages of parallel processing and phased processing, and can also reduce the number of read and write times of the memory unit 110, it further has the benefits of reducing power consumption and improving processing efficiency.
[0198] In summary, although the present invention has been disclosed by way of embodiments, it is not intended to limit the present invention. A person skilled in the art of the present invention may make various modifications and alterations without departing from the spirit and scope of the present invention. Therefore, the protection scope of the present invention shall be determined by the following claims.
Claims
1. An artificial intelligence algorithm computing acceleration processor, suitable for computing an input data in a memory unit, It is characterized in that The memory unit includes: a first data storage area for storing the input data; a second data storage area for storing a description data, wherein the description data includes a weight data; and a third data storage area for storing an output data; The artificial intelligence algorithm operation acceleration processor includes: a first temporary storage area for temporarily storing a first part of the input data, wherein the first temporary storage area has a predetermined data length; a second temporary storage area, for temporarily storing a first part of the description data; a third temporary storage area, used for temporarily storing a first part of the weight data; a first operator, for operating the first part of the input data and the first part of the weight data to generate a first operation result; a fourth temporary storage area, used for temporarily storing the first operation result; a fifth temporary storage area, for temporarily storing a second part of the weight data; and a second operator, used for operating the first operation result and the second part of the weight data to generate a second operation result; When the fourth temporary storage area is full of a predetermined amount of data, the second operator is triggered to operate the first operation result and the second part of the weight data.
2. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that When the second operator is triggered to start operating, the first operator continues to operate the input data.
3. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that The setting of the predetermined data amount for triggering the second operator is determined according to a batch width and a feature extractor parameter.
4. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that More selectively, the method includes: an activation unit for performing an activation operation on the first operation result.
5. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that More selectively, the method includes: a pooling unit for performing a pooling operation on the first operation result outputted from the fourth temporary storage area.
6. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that The first computing unit further includes a first computing unit array, the first computing unit array includes a plurality of first computing units, each of the first computing units is configured to: Receiving the input data and the first portion of the weight data corresponding to a multi-dimensional position; and The input data and the first part of the weight data are processed to generate a plurality of operation result values as the first operation result.
7. The artificial intelligence algorithm computing acceleration processor as claimed in claim 6, It is characterized in that The second computing unit further includes a second computing unit array, the second computing unit array includes a plurality of second computing units, each of the second computing units is configured to: receiving the first operation result and the second part of the weight data; and The first operation result and the second part of the weight data are processed to generate a plurality of operation result values as the second operation result.
8. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that The first operator has a first maximum operation amount, and the second operator has a second maximum operation amount, which is smaller than the first maximum operation amount.
9. The artificial intelligence algorithm computing acceleration processor as claimed in claim 1, It is characterized in that The fourth temporary storage area includes a capacity at least three times the predetermined data length.
10. The artificial intelligence algorithm computing acceleration processor as claimed in claim 7, It is characterized in that The number of the first operating units is greater than the number of the second operating units.
11. A method for accelerating computational processing of artificial intelligence algorithms. It is characterized in that The following steps are involved: A reads an input data and a description data from a memory unit, wherein the description data further includes a weight data; B. operating a first part of the input data and a first part of the weight data with a first operator to generate a first operation result; C temporarily stores the first operation result; D. When the first operation result stores a predetermined amount of data, triggering a second operation unit to operate the first operation result and a second part of the weight data by the second operation unit to generate a second operation result; as well as E writes the second operation result into the memory unit.
12. The artificial intelligence algorithm accelerated computing method of claim 11, It is characterized in that In the step D, when the second operator performs the second operation, the first operator and the second operator are in a parallel processing state.
13. The artificial intelligence algorithm accelerated computing method of claim 11, It is characterized in that The step A further comprises the following steps: A01 reads the first part of the input data from the memory unit to a first temporary storage area; A03 reads a first portion of the description data from the memory unit to a second temporary storage area; as well as A05 reads the first portion of the weight data from the memory unit to a third temporary storage area.
14. The artificial intelligence algorithm accelerated computing method of claim 13, It is characterized in that The step C further includes temporarily storing the first operation result of the first operator in a fourth temporary storage area.
15. The artificial intelligence algorithm accelerated computing method of claim 14, It is characterized in that The step A further comprises the following steps: A07 reads a second portion of the weight data from the memory unit to a fifth temporary storage area.
16. The artificial intelligence algorithm accelerated computing method of claim 15, It is characterized in that After step C, the method further comprises: F determines whether all the input data in the first temporary storage area have been read out and calculated. If the answer of step F is no, the next batch of input data is loaded from the first temporary storage area. If the answer of step F is yes, the process continues to step G; G determines whether all data in the fourth temporary storage area have been processed, if step G is no, then the data address parameter is updated, and if step G is yes, then the process continues to step H; and H determines whether there is input data in the first temporary storage area that has not been read out. If step H is no, the operation process ends. The setting of the predetermined data amount for triggering the second operator is determined according to a batch width and a feature extractor parameter.
17. The artificial intelligence algorithm accelerated computing method of claim 11, It is characterized in that After step E, the following steps are further included: I determines whether all the data in the fourth temporary storage area have completed the second operation. If not, continue to read the data in the fourth temporary storage area that have not completed the second operation. If yes, update a data address and end.
18. The artificial intelligence algorithm accelerated computing method of claim 17, It is characterized in that After step I, the method further includes: performing an activation operation on the first operation result.
19. The artificial intelligence algorithm accelerated computing method of claim 17, It is characterized in that After step I, the method further includes: performing a pooling operation on the first operation result.
20. The artificial intelligence algorithm accelerated computing method of claim 14, It is characterized in that The first temporary storage area has a predetermined data length, and the fourth temporary storage area has a capacity at least three times the predetermined data length.
21. A computer system, It is characterized in that include: A memory unit includes: a first data storage area for storing input data; a second data storage area for storing description data, wherein the description data includes weight data; and a third data storage area for storing output data; a memory read / write controller, coupled to the memory unit, to control reading and writing of the memory unit; and An artificial intelligence algorithm computing acceleration processor is coupled to the memory read / write controller, and the artificial intelligence algorithm computing acceleration processor includes: a first temporary storage area for temporarily storing a first part of the input data, wherein the first temporary storage area has a predetermined data length; a second temporary storage area, for temporarily storing a first part of the description data; a third temporary storage area, used for temporarily storing a first part of the weight data; a first operator, for operating the first part of the input data and the first part of the weight data to generate a first operation result; a fourth temporary storage area, used for temporarily storing the first operation result; a fifth temporary storage area, for temporarily storing a second part of the weight data; and a second operator, used for operating the first operation result and the second part of the weight data to generate a second operation result; When the fourth temporary storage area is full of a predetermined amount of data, the second operator is triggered to operate the first operation result and the second part of the weight data.
22. The computer system of claim 21, It is characterized in that When the second operator is triggered to start operating, the first operator continues to operate the input data.
23. A non-transitory computer-readable medium, It is characterized in that A program code is stored that can be read and executed by a computer. When the program code is executed, the computer performs the following actions: A reads an input data and a description data from a memory unit, wherein the description data further includes a weight data; B. operating a first part of the input data and a first part of the weight data with a first operator to generate a first operation result; C temporarily stores the first operation result; D. When the first operation result stores a predetermined amount of data, triggering a second operation unit to operate the first operation result and a second part of the weight data by the second operation unit to generate a second operation result; as well as E writes the second operation result into the memory unit.
Citation Information
Patent Citations
Neural network hardware accelerator
CN111915003A
Power efficient multiply-accumulate circuitry
US20210011971A1