A hardware accelerator
Through tensor movement, shape reshaping and switching sequence operations in hardware accelerators, the problem of inefficient tensor computing in the prior art is solved, the switching frequency with the CPU is reduced, and the AI computing efficiency is improved.
Patent Information
- Application Number
- CN202011196269.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-10-30
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2040-10-30
AI Technical Summary
Existing AI hardware accelerators are inefficient when performing tensor operations, and frequent switching with the CPU leads to low overall AI computing efficiency.
A hardware accelerator is designed, including memory, address generation unit and data queue. Through the combined operation of tensor movement, tensor reshaping shape and tensor swap order, the tensor depth to space arrangement or space to depth arrangement is realized, reducing frequent switching with the CPU.
It improves the efficiency of tensor computing, reduces the switching frequency between the hardware accelerator and the CPU, and improves the execution efficiency of the overall AI computing.
Smart Images

Figure CN114444675B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a spatial and depth conversion hardware accelerator for a tensor. Background Art
[0002] The explosive growth of artificial intelligence (AI) applications such as cloud big data processing and image analysis has prompted experts to continuously develop the best architectures. The goal of AI hardware accelerators is to find the fastest and most energy-efficient ways to perform the required computing tasks.
[0003] There are several major application scenarios for AI hardware acceleration applications: (1) Cloud acceleration involves compression and decompression, blockchain, and security, etc., and requires a high computing power and power consumption cost ratio; (2) Storage, some applications require high efficiency, so data processing in the memory is required; (3) Autonomous driving involves artificial intelligence, data computing, and sensor fusion, etc., and requires programmability.
[0004] With the emergence of AI hardware accelerators, how to make the hardware execute more efficiently in response to different AI neural network computing requirements is one of the efforts in the industry.
[0005] In addition, how to reduce the frequent switching problem between the AI hardware accelerator and the CPU and reduce the CPU load can also take into account the overall AI operation execution efficiency. Summary of the Invention
[0006] According to an embodiment of the present invention, a hardware accelerator is provided, which includes: a first memory for receiving data; a source address generation unit coupled to the first memory for generating a plurality of source addresses for the first memory according to a source shape parameter and a source movement parameter, such that the first memory sends out data according to these source addresses; a data collection unit coupled to the first memory for receiving the data transmitted from the first memory; a first data queue coupled to the data collection unit for temporarily storing the data transmitted from the data collection unit; a data dispersion unit coupled to the first data queue for dispersing the data transmitted from the first data queue; a target address generation unit coupled to the data dispersion unit for generating a plurality of target addresses according to a target shape parameter and a target movement parameter; an address queue coupled to the target address generation unit for temporarily storing the target addresses generated by the target address generation unit; a second data queue coupled to the data dispersion unit for temporarily storing the data transmitted from the data dispersion unit; and a second memory coupled to the second data queue for writing the data transmitted from the second data queue according to the target addresses generated by the target address generation unit. The hardware accelerator performs any combination of tensor movement, tensor reshaping, and tensor transposing to achieve tensor depth-to-space permutation or tensor space-to-depth permutation.
[0007] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments, but is not limited to the present invention. Description of the Drawings
[0008] Figure 1A Schematic diagram showing 4D tensor data stored in a memory.
[0009] Figure 1B Another example showing 4D tensor data stored in a memory.
[0010] Figure 2 Functional block diagram showing a hardware accelerator according to an embodiment of the present invention.
[0011] Figure 3 An example showing "tensor depth-to-space permutation" and "tensor space-to-depth permutation".
[0012] Figure 4 Flowchart showing source address generation according to an embodiment of the present invention.
[0013] Figure 5 Flowchart showing target address generation according to an embodiment of the present invention.
[0014] Figure 6A and Figure 6BShow two exemplary tensor shifts according to an embodiment of the present invention.
[0015] Figure 7 Show an exemplary tensor reshaping according to an embodiment of the present invention.
[0016] Figure 8 Show an exemplary tensor transposition according to an embodiment of the present invention.
[0017] Figures 9A to 9C Show the tensor depth-to-space arrangement algorithm according to an embodiment of the present invention. Detailed Description of the Invention
[0018] The structural principle and working principle of the present invention will be specifically described below with reference to the accompanying drawings:
[0019] In deep learning, a tensor is actually a multidimensional array. According to its dimension, tensors can be divided into 1D tensors, 2D tensors, 3D tensors, 4D tensors, etc. In the present invention, the data is described as a 4D tensor, but it should be understood that the present invention is not limited thereto. In other possible embodiments of the present invention, the data can be a 3D tensor or a higher-dimensional tensor (such as a 5D tensor), etc., which are all within the spirit of the present invention.
[0020] Figure 1A Show a schematic diagram of storing 4D tensor data in the memory 100. In the present invention, the shape parameter of the 4D tensor data is defined as: (n, c, h, w), which represents the number of points in the four-dimensional shape; and the stride parameter of the 4D tensor data is defined as: (nStride, cStride, hStride, wStride), which represents the distance from point to point. That is, nStride (which can also be abbreviated as ns) represents the distance from point n to point n, cStride (which can also be abbreviated as cs) represents the distance from point c to point c, hStride (which can also be abbreviated as hs) represents the distance from point h to point h, and wStride (which can also be abbreviated as ws) represents the distance from point w to point w. BA represents the base address, so the base address can also be abbreviated as BA.
[0021] For easier understanding, Figure 1BAnother exemplary illustration of storing 4D tensor data in memory 100. The shape parameters of this 4D tensor data are (n, c, h, w) = (2, 3, 4, 4), and the movement parameters of the 4D tensor data are defined as (ns, cs, hs, ws) = (48, 16, 4, 1). Among them, the image (n, c, h, w) = (2, 3, 4, 4) means there are 2 frames, each frame has 3 channels (R, G, B), each R plane, G plane, and B plane is composed of 4*4 h*w points respectively, with a total of 4*4*3*2 = 96 points, and each point is 1 byte.
[0022] Figure 2 A functional block diagram of a hardware accelerator according to an embodiment of the present invention is shown. The hardware accelerator 200 includes: a first memory 210, a data collection unit 215, a source address generation unit 220, a first data queue 225, a data dispersion unit 230, a target address generation unit 235, an address queue 240, a second data queue 245, and a second memory 250. In possible embodiments of the present invention, the first memory 210 and the second memory 250 can be: (1) implemented by the same memory unit, respectively occupying different addresses of the memory unit; or, (2) implemented by the same memory unit, using an in-place algorithm, respectively occupying the same address of the memory unit; or, (3) implemented by different memory units.
[0023] The hardware accelerator 200 of this embodiment can perform tensor stride, tensor reshape, and tensor transpose. Moreover, the hardware accelerator 200 of this embodiment can perform any combination of tensor stride, tensor reshape, and tensor transpose to achieve tensorflow depth-to-space permutation and / or tensorflow space-to-depth permutation. Tensorflow is an operation term defined in a certain AI framework. As for other AI frameworks (such as pytorch, etc.), there may be different names, which are all within the spirit of the present invention.
[0024] The first memory 210 receives data transmitted from a data source (not shown in the figure).
[0025] The data collection unit 215 is coupled to the first memory 210 and receives the data transmitted from the first memory 210.
[0026] The source address generation unit 220 is coupled to the data collection unit 215 to generate a plurality of source addresses SAD (source address) according to the source shape parameters (n, c, h, w) and the source movement parameters (ns, cs, hs, ws). The source shape parameters (n, c, h, w) and the source movement parameters (ns, cs, hs, ws) are user-defined.
[0027] The source address generation unit 220 sends the source address SAD to the first memory 210, and the first memory 210 sends data to the data collection unit 215 according to the source address SAD. The data collection unit 215 collects the data sent by the first memory 210.
[0028] The first data queue 225 is coupled to the data collection unit 215 to temporarily store the data transmitted from the data collection unit 215.
[0029] The data dispersion unit 230 is coupled to the first data queue 225 to disperse the data transmitted from the first data queue 225. Among them, the data transmitted from the first data queue 225 is the data written into the first memory 210 according to the source shape parameters (n, c, h, w) and the source movement parameters (ns, cs, hs, ws).
[0030] The destination address generation unit 235 is coupled to the data dispersion unit 230 to generate a plurality of destination addresses DAD (Destination address) according to the destination shape parameters (n', c', h', w') and the destination movement parameters (ns', cs', hs', ws'). The destination shape parameters (n', c', h', w') and the destination movement parameters (ns', cs', hs', ws') are user-defined.
[0031] The address queue 240 is coupled to the destination address generation unit 235 to temporarily store the destination addresses DAD generated by the destination address generation unit 235.
[0032] The second data queue 245 is coupled to the data dispersion unit 230 to temporarily store the data transmitted from the data dispersion unit 230.
[0033] The second memory 250 is coupled to the second data queue 245 to write the data transmitted from the second data queue 245 according to the destination addresses DAD generated by the destination address generation unit 235.
[0034] Figure 3Shows an example of "tensor depth-to-space permutation" and "tensor space-to-depth permutation". Tensor depth-to-space permutation can rearrange data from depth data into spatial data; and, tensor space-to-depth permutation can rearrange data from spatial data into depth data. Figure 3 Above represents depth data, and Figure 3 below represents spatial data. Taking Figure 3 as an example, the shape parameters of this 4D tensor depth data are (n, cr 2 , h, w) = (1, 8, 3, 3), and the shape parameters of this 4D tensor spatial data are (n, c, rh, rw) = (1, 2, 6, 6). In Figure 3 , the parameter r = 2 represents the block size.
[0035] Figure 4 Shows a flowchart for generating a source address according to an embodiment of the present invention. At the beginning, the parameters i, j, k, and l are all set to 0. In Figure 4 , if any of the parameters i, j, k, and l is updated, then a new source address SAD is generated, where SAD = BA + l * ns + k * cs + j * hs + i * ws. Here, BA represents the base address (i.e., the starting address) of the source address.
[0036] In step 405, perform a W permutation of the data (i.e., read the data in the W direction from the first memory 210). Then, in step 407, increment i by 1. After that, in step 410, determine whether the parameter i is less than the parameter w. If so, the process returns to step 405 and generates the source address SAD. If step 410 is false, then in step 412, set the parameter i = 0 and the process continues to step 415 to perform an H permutation of the data (i.e., read the data in the H direction from the first memory 210). Then, in step 417, increment j by 1.
[0037] In step 420, determine whether the parameter j is less than the parameter h. If so, the process returns to step 405 and generates the source address SAD. If step 420 is false, then in step 422, set the parameter j = 0 and the process continues to step 425 to perform a C permutation of the data (i.e., read the data in the C direction from the first memory 210).
[0038] Next, in step 427, increment k by 1. In step 430, determine whether the parameter k is less than the parameter c. If so, the process returns to step 405 and generates the source address SAD. If step 430 is false, then in step 432, set the parameter k = 0 and the process continues to step 435 to perform an N permutation of the data (i.e., read the data in the N direction from the first memory 210).
[0039] Next, in step 437, increment parameter l by 1. In step 440, determine whether parameter l is less than parameter n. If so, the process returns to step 405 and generates the source address SAD. If step 440 is false, then in step 442, set parameter l = 0 and end the process.
[0040] As can be seen from the above process, there are a total of n*c*h*w source addresses SAD generated.
[0041] Figure 5 Show the flowchart for generating the target address according to an embodiment of the present invention. At the beginning, set parameters i, j, k, and l to 0. In Figure 5 If any of the parameters i, j, k, and l is updated, then a new target address DAD is generated, where DAD = BA'+l*ns'+k*cs'+j*hs'+i*ws'. Here, BA' represents the base address (i.e., the starting address) of the target address.
[0042] In step 505, perform the W arrangement of the data (i.e., write the data into the second memory 250 in the W direction). Next, in step 507, increment i by 1. Then, in step 510, determine whether parameter i is less than parameter w'. If so, the process returns to step 505 and generates the target address DAD. If step 510 is false, then in step 512, set parameter i = 0 and the process continues to step 515 to perform the H arrangement of the data (i.e., write the data into the second memory 250 in the H direction). Next, in step 517, increment j by 1.
[0043] In step 520, determine whether parameter j is less than parameter h'. If so, the process returns to step 505 and generates the target address DAD. If step 520 is false, then in step 522, set parameter j = 0 and the process continues to step 525 to perform the C arrangement of the data (i.e., write the data into the second memory 250 in the C direction).
[0044] Next, in step 527, increment k by 1. In step 530, determine whether parameter k is less than parameter c'. If so, the process returns to step 505 and generates the target address DAD. If step 530 is false, then in step 532, set parameter k = 0 and the process continues to step 535 to perform the N arrangement of the data (i.e., write the data into the second memory 250 in the N direction).
[0045] Next, in step 537, increment the parameter l by 1. In step 540, determine whether the parameter l is less than the parameter n'. If so, the process returns to step 505 and the target address DAD is generated. If step 540 is false, then in step 542, set the parameter l = 0 and end the process.
[0046] As can be seen from the above process, the total number of generated target addresses DAD is n' * c' * h' * w'.
[0047] Now, several ways of permutation in the embodiments of the present invention will be described.
[0048] (A) Tensor copy with stride
[0049] Tensor copy with stride can be applied to "tensor depth-to-space permutation" and "tensor space-to-depth permutation".
[0050] Figure 6A And Figure 6B shows two exemplary tensor copy with stride according to an embodiment of the present invention. In Figure 6A , the source shape parameters (n, c, h, w) = (2, 2, 2, 3), and the source stride parameters (ns, cs, hs, ws) = (12, 6, 3, 1). The target shape parameters (n', c', h', w') = (2, 2, 2, 3), and the target stride parameters (ns', cs', hs', ws') = (16, 8, 4, 1).
[0051] In Figure 6B , the source shape parameters (n, c, h, w) = (2, 2, 2, 3), and the source stride parameters (ns, cs, hs, ws) = (12, 6, 3, 1). The target shape parameters (n', c', h', w') = (2, 2, 2, 3), and the target stride parameters (ns', cs', hs', ws') = (8, 16, 4, 1).
[0052] That is to say, after tensor copy with stride, the shapes of the source data and the target data are the same, but the positions where they are placed are different.
[0053] (B) Tensor reshape
[0054] Tensor reshape can be applied to "tensor depth-to-space permutation" and "tensor space-to-depth permutation".
[0055] Figure 7 shows an exemplary tensor reshape according to an embodiment of the present invention. In Figure 7Among them, the source shape parameters (n, c, h, w) = (2, 1, 2, 3), and the source movement parameters (ns, cs, hs, ws) = (16, 8, 4, 1). The target shape parameters (n', c', h', w') = (2, 1, 3, 2), and the target movement parameters (ns', cs', hs', ws') = (16, 8, 4, 1).
[0056] That is to say, after reshaping the tensor, the shapes of the source data and the target data are different (but the product of the total number of points is the same).
[0057] (C) Tensor transpose
[0058] Tensor transpose can be applied to "tensor depth-to-space permutation" and "tensor space-to-depth permutation".
[0059] Figure 8 Show a demonstration example of tensor transpose according to an embodiment of the present invention. In Figure 8 Among them, the source shape parameters (n, c, h, w) = (1, 2, 1, 3), and the source movement parameters (ns, cs, hs, ws) = (16, 8, 4, 1). The target shape parameters (n', c', h', w') = (1, 3, 1, 2), and the target movement parameters (ns', cs', hs', ws') = (16, 8, 4, 1).
[0060] In Figure 8 Among them, it can be regarded as swapping the parameters c and w. In other possible examples of the present invention, any two shape parameters can be swapped (such as swapping n and w... etc.), which are all within the spirit of the present invention.
[0061] That is to say, after performing tensor transpose, the shapes of the source data and the target data are different (because two shape parameters are swapped, resulting in different shapes).
[0062] Now, how the embodiment of the present invention achieves "tensor depth-to-space permutation" will be described.
[0063] The first tensor depth-to-space permutation algorithm:
[0064] Taking the rearrangement of the source shape parameters (n, cr 2 , h, w) = (1, 2 * 2 2 , 3, 3) = (1, 8, 3, 3) into the target shape parameters (n', c', rh', rw') = (1, 2, 2 * 3, 2 * 3) = (1, 2, 6, 6) as an example for illustration.
[0065] The first tensor depth-to-space permutation algorithm includes three steps.
[0066] In the first step, the order of tensors is swapped to swap parameter c with w. After the operation, the source shape parameter (n, c, h, w) = (1, 8, 3, 3) and the source movement parameter (ns, cs, hs, ws) = (72, 9, 3, 1). After swapping the order of tensors to swap parameter c with w, the target shape parameter (n', c', h', w') = (1, 3, 3, 8) and the target movement parameter (ns', cs', hs', ws') = (72, 24, 8, 1), as Figure 9A shown.
[0067] In the second step, the result of the first step is reshaped and moved in tensors.
[0068] As Figure 9B shown, after reshaping and moving the source shape parameter (n, c, h, w) = (1, 3, 3, 8) and the source movement parameter (ns, cs, hs, ws) = (72, 24, 8, 1) in tensors, the target shape parameter (n', c', h', w') = (3, 6, 2, 2) and the target movement parameter (ns', cs', hs', ws') = (2(=r), 12(=r*r*w), 6(=r*w), 1) can be obtained.
[0069] Next, when moving the tensors in the second step, considering n*h*w as the new w, after moving the tensors, we can get: the target shape parameter (n', c', h', w') = (1, 6, 1, 12) and the target movement parameter (ns', cs', hs', ws') = (-, 12, -, 1), where the symbol "-" represents don't care.
[0070] In the third step, the result of the second step is reshaped and moved in tensors.
[0071] As Figure 9C shown, after reshaping and moving the source shape parameter (n, c, h, w) = (1, 6, 1, 12) and the source movement parameter (ns, cs, hs, ws) = (-, 12, -, 1) in tensors, the target shape parameter (n', c', h', w') = (1, 3, 2, 12) and the target movement parameter (ns', cs', hs', ws') = (-, 12, 36, 1) can be obtained.
[0072] In the first tensor depth-to-space arrangement algorithm, "tensor depth-to-space arrangement" can be achieved by (1) swapping the order of tensors; (2) reshaping and moving the tensors; and (3) reshaping and moving the tensors.
[0073] The second tensor depth-to-space permutation algorithm:
[0074] Rearrange the source shape parameters (n, c, h, w) = (1, 2048, 8, 6) into the target shape parameters (n', c', h', w') = (1, 512, 16, 12). Here, define the parameter c as c = c1 * r 2 , where r represents the block size. For the above example, c = 2048, c1 = 512, and r = 2.
[0075] The second tensor depth-to-space permutation algorithm includes tensor moves repeated multiple times.
[0076] In the first tensor move, when performing, substitute the parameter r 2 into the parameter c, that is, the source shape parameters (n, c, h, w) = (1, 4(= r 2 ), 8, 6). After performing, the result is as follows: the source shape parameters (n, c, h, w) = (1, 4(= r 2 ), 8, 6) and the source move parameters (ns, cs, hs, ws) = (-, 48, 6, 1). After performing the tensor move, the target shape parameters (n', c', h', w') = (2, 2, 8, 6) and the target move parameters (ns', cs', hs', ws') = (12(= r * w), 1, 24(= r * r * w), 2(= r)).
[0077] After that, repeat the above tensor move c1 times.
[0078] That is to say, in the second tensor depth-to-space permutation algorithm, perform the operation on one c1 plane at a time, so it needs to be repeated multiple times.
[0079] In the second tensor depth-to-space permutation algorithm, through repeated tensor moves, "tensor depth-to-space permutation" can be achieved.
[0080] The third tensor depth-to-space permutation algorithm:
[0081] Rearrange the source shape parameters (n, c, h, w) = (1, 2048, 8, 6) into the target shape parameters (n', c', h', w') = (1, 512, 16, 12). Here, define the parameter c as c = c1 * r 2 , where r represents the block size. For the above example, c = 2048, c1 = 512, and r = 2.
[0082] The third tensor depth-to-space permutation algorithm includes repeating the tensor move step 2 r times.
[0083] In the first tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6). After performing, the results are as follows: the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddr. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 12(=r*w), 2(=r)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddr.
[0084] In the second tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6). After performing, the results are as follows: the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddr + 1*h*w = SourceBaseAddr + 48. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 12(=r*w), 2(=r)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddr + 1.
[0085] In the 3rd tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6). After performing, the results are as follows: the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddr + 2*h*w = SourceBaseAddr + 96. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 12(=r*w), 2(=r)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddr + r*w = DestinationBaseAddr + 12.
[0086] In the 4th tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6). After performing, the results are as follows: the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddr + 3*h*w = SourceBaseAddr + 144. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 12(=r*w), 2(=r)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddr + r*w + 1 = DestinationBaseAddr + 13.
[0087] When performing multiple tensor movements, the base address BA of the source address increases by h*w each time.
[0088] When performing multiple tensor movements, the base address BA' of the target address is DestinationBaseAddress + offset1. offset1 = (i - 1)*r*w, (i - 1)*r*w + 1, …, (i - 1)*r*w + (i - 1), where i = 1 ~ r.
[0089] In the third tensor depth-to-space permutation algorithm, "tensor depth-to-space permutation" can be achieved by performing multiple tensor moves.
[0090] Now, it will be described how the embodiments of the present invention achieve "tensor space-to-depth permutation". Basically, the "tensor space-to-depth permutation" of the embodiments of the present invention can be regarded as the reverse operation of the "tensor depth-to-space permutation" of the above embodiments of the present invention.
[0091] The first tensor space-to-depth permutation algorithm:
[0092] Taking the rearrangement of the source shape parameters (n, c, h, w) = (1, 2, 6, 6) into the target shape parameters (n', c', h', w') = (1, 8, 3, 3) as an example for illustration. Here, the parameter c' is defined as c = c1 * r 2 , where r represents the block size. For the above example, c' = 8, c1 = 2, and r = 2.
[0093] The first tensor space-to-depth permutation algorithm includes three steps.
[0094] In the first step, tensor reshaping and tensor moving are performed.
[0095] From the source shape parameters (n, c, h, w) = (1, 2, 6, 6), the target shape parameters (n', c', h', w') = (1, 2, 3, 12) are obtained. In this step, no moving is performed, but it is regarded as a tightly packed shape.
[0096] After that, after performing tensor moving on the source shape parameters (n, c, h, w) = (1, 2, 3, 12), the target shape parameters (n', c', h', w') = (1, 2, 3, 12) and the target moving parameters (ns', cs', hs', ws') = (-, 12 (= r * r * w), 24 (= r * r * w * c1), 1) can be obtained.
[0097] In the second step, the result of the first step is subjected to tensor reshaping and tensor moving.
[0098] From the source shape parameters (n, c, h, w) = (1, 2, 3, 12), the target shape parameters (n', c', h', w') = (3, 4, 3, 2) are obtained. In this step, no moving is performed, but it is regarded as a tightly packed shape.
[0099] Next, after performing the second step of tensor shifting on the source shape parameters (n, c, h, w) = (3, 4, 3, 2), the following can be obtained: the target shape parameters (n', c', h', w') = (3, 4, 3, 2) and the target shifting parameters (ns', cs', hs', ws') = (8(=c1*r*r), 2(=r), 24(=r*r*c1*h), 1).
[0100] In the third step, tensor reshaping and tensor permutation order are performed. From the source shape parameters (n, c, h, w) = (3, 4, 3, 2), the target shape parameters (n', c', h', w') = (1, 3, 3, 8) are obtained. In this step, no shifting is performed, but it is regarded as a tightly packed shape.
[0101] Next, the tensor permutation order is performed on the source shape parameters (n, c, h, w) = (1, 3, 3, 8) to swap the parameters c and w, and the target shape parameters (n', c', h', w') = (1, 8, 3, 3) are obtained.
[0102] In the first tensor space-to-depth arrangement algorithm, "tensor space-to-depth arrangement" can be achieved through (1) performing tensor reshaping and tensor shifting; (2) performing tensor reshaping and tensor shifting; and (3) performing tensor reshaping and tensor permutation order.
[0103] The second tensor space-to-depth arrangement algorithm:
[0104] Taking the rearrangement of the source shape parameters (n, c, h, w) = (1, 2, 6, 6) into the target shape parameters (n', c', h', w') = (1, 8, 3, 3) as an example for illustration. Here, the parameter c' is defined as c = c1*r 2 , where r represents the block size. In the above example, c' = 8, c1 = 2, and r = 2.
[0105] The second tensor space-to-depth arrangement algorithm includes tensor shifting repeated multiple times (i.e., c1 times).
[0106] In the first tensor shift, the source shape parameters (n, c, h, w) = (3, 2, 3, 2), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress. After performing the tensor shift, the target shape parameters (n', c', h', w') = (3, 2, 3, 2) and the target shifting parameters (ns', cs', hs', ws') = (3(=w), 18(=r*h*w), 1, 9(=w*h)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress.
[0107] In the second tensor shift, the source shape parameters (n, c, h, w) = (3, 2, 3, 2), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress + r * r * w * h = SourceBaseAddress + 36. After the tensor shift, the target shape parameters (n', c', h', w') = (3, 2, 3, 2) and the target shift parameters (ns', cs', hs', ws') = (3 (= w), 18 (= r * h * w), 1, 9 (= w * h)), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress + r * r * w * h = DestinationBaseAddress + 36.
[0108] In the second tensor space-to-depth permutation algorithm, "tensor space-to-depth permutation" can be achieved by repeatedly performing multiple tensor shifts.
[0109] The third tensor space-to-depth permutation algorithm:
[0110] Rearrange the source shape parameters (n, c, h, w) = (1, 512, 16, 12) into the target shape parameters (n', c', h', w') = (1, 2048, 8, 6). Here, define the parameter c as c = c1 * r 2 , where r represents the block size. Taking the above example, c = 2048, and c1 = 512, r = 2.
[0111] The third tensor space-to-depth permutation algorithm includes repeatedly performing r 2 times of tensor shift steps.
[0112] In the first tensor shift step, when performing, substitute the parameter c1 into the parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512 (= c1), 8, 6), the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source shift parameters (ns, cs, hs, ws) = (-, 192 (= r * r * h * w), 24 (= r * r * w), 2 (= r)), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress. After the tensor shift, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target shift parameters (ns', cs', hs', ws') = (-, 192 (= r * r * h * w), 6 (= w), 1), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress.
[0113] In the second tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6), the source shape parameters (n, c, h, w) = (1, 512, 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 24(=r*r*w), 2(=r)), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress + 1. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress + 1*h*w = DestinationBaseAddress + 48.
[0114] In the third tensor movement step, when performing, substitute parameter c1 into parameter c, that is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6) and the source movement parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 24(=r*r*w), 2(=r)), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress + r*w = SourceBaseAddress + 12. After performing the tensor movement, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target movement parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress + 2*h*w = DestinationBaseAddress + 96.
[0115] In the 4th tensor shift step, when performing, substitute parameter c1 into parameter c. That is, the source shape parameters (n, c, h, w) = (1, 512(=c1), 8, 6) and the source shift parameters (ns, cs, hs, ws) = (-, 192(=r*r*h*w), 24(=r*r*w), 2(=r)), and the base address BA (i.e., the starting address) of the source address = SourceBaseAddress + r*w + 1 = SourceBaseAddress + 13. After performing the tensor shift, the target shape parameters (n', c', h', w') = (1, 512, 8, 6) and the target shift parameters (ns', cs', hs', ws') = (-, 192(=r*r*h*w), 6(=w), 1), and the base address BA' (i.e., the starting address) of the target address = DestinationBaseAddress + 3*h*w = DestinationBaseAddress + 144.
[0116] When performing multiple tensor shifts, the base address BA of the source address is SourceBaseAddress + offset2. offset2 = (i - 1)*r*w, (i - 1)*r*w + 1, …, (i - 1)*r*w + (i - 1), where i = 1 ~ r.
[0117] When performing multiple tensor shifts, the base address BA' of the target address increases by h*w each time.
[0118] In the third tensor space to depth permutation algorithm, by performing multiple tensor shifts, "tensor space to depth permutation" can be achieved.
[0119] From the above, it can be seen that the embodiments of the present invention can achieve "tensor space to depth permutation" and "tensor depth to space permutation" more efficiently, and can reduce the frequent switching between the hardware accelerator and the central processing unit (CPU), reduce the CPU load, and increase the overall AI operation execution efficiency.
[0120] Of course, the present invention may have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A hardware accelerator, characterized in that, Comprising: A first memory for receiving data; A source address generation unit coupled to the first memory for generating a plurality of source addresses for the first memory according to a source shape parameter and a source movement parameter, such that the first memory sends out data according to the source addresses; A data collection unit coupled to the first memory for receiving the data transmitted from the first memory; A first data queue coupled to the data collection unit for temporarily storing the data transmitted from the data collection unit; A data dispersion unit coupled to the first data queue for dispersing the data transmitted from the first data queue; A target address generation unit coupled to the data dispersion unit for generating a plurality of target addresses according to a target shape parameter and a target movement parameter; An address queue coupled to the target address generation unit for temporarily storing the target addresses generated by the target address generation unit; A second data queue coupled to the data dispersion unit for temporarily storing the data transmitted from the data dispersion unit; And A second memory coupled to the second data queue for writing the data transmitted from the second data queue according to the target addresses generated by the target address generation unit, wherein the hardware accelerator can perform one or any combination of tensor movement, tensor reshaping, and tensor permutation to achieve tensor depth-to-space arrangement and / or tensor space-to-depth arrangement; wherein, by performing tensor permutation, performing tensor reshaping and tensor movement, and performing tensor reshaping and tensor movement, or by performing multiple tensor movements, tensor depth-to-space arrangement is achieved.
2. A hardware accelerator, characterized in that, Comprising: A first memory for receiving data; A source address generation unit coupled to the first memory for generating a plurality of source addresses for the first memory according to a source shape parameter and a source movement parameter, such that the first memory sends out data according to the source addresses; A data collection unit coupled to the first memory for receiving the data transmitted from the first memory; A first data queue coupled to the data collection unit for temporarily storing the data transmitted from the data collection unit; A data dispersion unit coupled to the first data queue for dispersing the data transmitted from the first data queue; A target address generation unit coupled to the data dispersion unit for generating a plurality of target addresses according to a target shape parameter and a target movement parameter; An address queue coupled to the target address generation unit for temporarily storing the target addresses generated by the target address generation unit; A second data queue coupled to the data dispersion unit for temporarily storing the data transmitted from the data dispersion unit; And A second memory coupled to the second data queue for writing the data transmitted from the second data queue according to the target addresses generated by the target address generation unit, Among them, the hardware accelerator can perform one or any combination of tensor movement, tensor reshaping, and tensor transposition to achieve tensor depth-to-space arrangement and / or tensor space-to-depth arrangement; Among them, by performing tensor reshaping and tensor movement, performing tensor reshaping and tensor movement, and performing tensor reshaping and tensor transposition, or by performing multiple tensor movements, to achieve tensor space-to-depth arrangement.
3. The hardware accelerator according to claim 1 or 2, characterized in that Among them, The source shape parameter and the source movement parameter are defined by the user.
4. The hardware accelerator according to claim 1 or 2, characterized in that, Among them, The target shape parameter and the target movement parameter are defined by the user.
5. The hardware accelerator according to claim 1 or 2, characterized in that Among them, After performing tensor movement, the source shape parameter is the same as the target shape parameter, while the source movement parameter is different from the target movement parameter.
6. The hardware accelerator according to claim 1 or 2, characterized in that, Among them, After performing tensor reshaping, the source shape parameter is different from the target shape parameter.
7. The hardware accelerator according to claim 1 or 2, characterized in that Among them, After performing tensor transposition, the source shape parameter is different from the target shape parameter.
Citation Information
Patent Citations
Data processing method and device, computer equipment and storage medium
CN111401539A
Computer-implemented tensor data calculation method and device, medium and equipment
CN111767508A
Memory operand descriptors
US20170262174A1