Processing device, processing method, electronic device, medium, program product
By adopting multiple computing units and memory in the processing device and using the data replication mechanism between the computing units, the problem of slow data loading speed in tensor data calculation in the prior art is solved, and more efficient data loading and computing performance is achieved.
Patent Information
- Application Number
- CN202410487670.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-04-22
AI Technical Summary
When the existing graphics processing unit performs tensor data calculation, the calculation unit needs to frequently load data from external memory, resulting in frequent memory access operations and slow loading speed, which affects the computing performance.
By designing multiple computing units and corresponding memory in the processing device, and adopting a data copying mechanism between the computing units, the amount and time of each computing unit loading data from the external memory is reduced. The specific method is that the first computing unit loads part of the segment of the tensor into its corresponding memory, and copies the segment to the memory of other computing units to realize data sharing and reduce duplicate loading.
By reducing the dependence of the computing unit on external memory, the data loading speed and computing performance are improved, and since operations between the computing units can be performed in parallel, the data preparation and computing process are further accelerated.
Smart Images

Figure CN118297788B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of processing devices, and more particularly, to processing devices, processing methods, electronic devices, non-transitory computer-readable storage media, and computer program products in the field of artificial intelligence. Background Art
[0002] A Graphics Processing Unit (GPU) typically uses multiple computing units in a streaming processor cluster to perform operations on tensor data, such as operations of neural network models and matrix multiplications. How to fully improve the computational efficiency of operations is an issue that needs to be considered. Summary of the Invention
[0003] According to one aspect of the present application, there is provided a processing device, including: a plurality of computing units and a plurality of memories corresponding to the plurality of computing units; wherein, a first computing unit among the plurality of computing units loads a first section of a first tensor part of a first tensor into a first memory corresponding to the first computing unit, and copies the first section of the first tensor part to a second memory corresponding to a second computing unit among the plurality of computing units.
[0004] According to another aspect of the present application, there is provided a processing method for a processing device, wherein the processing device includes a plurality of computing units and a plurality of memories corresponding to the plurality of computing units, and the method includes: loading, by a first computing unit among the plurality of computing units, a first section of a first tensor part of a first tensor into a first memory corresponding to the first computing unit; and copying, by the first computing unit, the first section of the first tensor part to a second memory corresponding to a second computing unit among the plurality of computing units.
[0005] According to another aspect of the present application, there is provided an electronic device, including means for performing each step of the method according to an embodiment of the present application.
[0006] According to another aspect of the present application, there is provided an electronic device, including: a memory for storing computer instructions; and a processor for reading the computer instructions in the memory and executing the method according to an embodiment of the present application.
[0007] According to another aspect of the present application, there is provided a non-transitory computer-readable storage medium, on which computer instructions are stored, wherein, when the computer instructions are executed by a processor, the processor is caused to execute the method according to an embodiment of the present application.
[0008] According to another aspect of the present application, there is provided a computer program product including computer instructions, wherein when the computer instructions are executed by a processor, the processor is caused to execute the method according to the embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present disclosure, and for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0010] Figure 1 The block diagram of a processing device and the schematic diagram of an application scenario according to the embodiments of the present application are shown.
[0011] Figure 2 The schematic diagram of the processing operations of the processing device according to the embodiments of the present application is shown.
[0012] Figure 3 The schematic diagram of the operation of an array of 2 rows * 4 columns of computing units according to the embodiments of the present application is shown.
[0013] Figure 4A The flowchart of the processing method for a processing device according to the embodiments of the present application is shown.
[0014] Figure 4B The flowchart of the processing method for a processing device according to another embodiment of the present application is shown.
[0015] Figure 4C The flowchart of the processing method for a processing device according to another embodiment of the present application is shown.
[0016] Figure 5 The flowchart of the processing method for a processing device according to another embodiment of the present application is shown.
[0017] Figure 6 The flowchart of the synchronization mechanism for a memory to receive data and a memory to consume data included in the processing method according to the embodiments of the present application is shown.
[0018] Figure 7 The block diagram of an electronic device according to the embodiments of the present application is shown.
[0019] Figure 8 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application is shown.
[0020] Figure 9Shows a schematic diagram of a non - transitory computer - readable storage medium according to an embodiment of the present application. Detailed implementation
[0021] Now, specific embodiments of the present application will be described in detail. Examples of the present application are illustrated in the accompanying drawings. Although the present application will be described in conjunction with specific embodiments, it will be understood that it is not intended to limit the present application to the described embodiments. On the contrary, it is intended to cover modifications, variations, and equivalents included within the spirit and scope of the present application. It should be noted that the method steps described herein can all be implemented by any functional block or functional arrangement, and any functional block or functional arrangement can be implemented as a physical entity or a logical entity, or a combination of both.
[0022] Figure 1 Shows a block diagram of a processing device and a schematic diagram of an application scenario according to an embodiment of the present application.
[0023] As Figure 1 shown, an operation is to be performed on a first tensor A and a second tensor B, such as a matrix multiplication or a convolution operation, etc. If the operation includes a matrix multiplication operation, the first tensor can be the left matrix, and the second tensor can be the right matrix. If the operation includes a convolution operation, the first tensor can be the input tensor, and the second tensor can be the weight tensor.
[0024] The processing device 100 includes: a plurality of calculation units (Calculation Unit, CU) (such as calculation units CU0, CU1, CU2, CU3) and a plurality of memories corresponding to the plurality of calculation units (such as memories M0, M1, M2, M3). Of course, the number of calculation units and memories is only an example, not a limitation.
[0025] These memories are equivalent to the internal memories of the calculation units. Among them, calculation unit CU0 corresponds to memory M0, calculation unit CU1 corresponds to memory M1, calculation unit CU2 corresponds to memory M2, and calculation unit CU3 corresponds to memory M3.
[0026] Multiple calculation units are needed to perform parallel calculations to implement the operation on the first tensor A and the second tensor B. Assume as Figure 1As shown, there are 4 computing units. The first tensor A is split into 2 first tensor parts A0 and A1, and the second tensor B is split into 2 second tensor parts B0 and B1. Then, in the prior art, generally, the computing unit CU0 needs to load the first tensor part A0 and the second tensor part B0 into its internal memory M0 and perform operations, such as multiplying and adding each number in A0 with each number in B0 one by one. The computing unit CU1 needs to load the first tensor part A0 and the second tensor part B1 into its internal memory M1 and perform operations. The computing unit CU2 needs to load the first tensor part A1 and the second tensor part B0 into its internal memory M2 and perform operations. The computing unit CU3 needs to load the first tensor part A1 and the second tensor part B1 into its internal memory M3 and perform operations. That is to say, the computing units CU0 and CU1 need to load the complete first tensor part A0 from the outside twice in total. The computing units CU0 and CU2 need to load the complete second tensor part B0 twice in total. The computing units CU2 and CU3 need to load the complete first tensor part A1 from the outside twice in total. The computing units CU1 and CU3 need to load the complete second tensor part B1 from the outside twice in total.
[0027] However, since each computing unit has to load all data from the outside, such as from the global memory, into its internal memory every time, a large number of memory access operations to the global memory are required, and the loading speed is slow, thus affecting the operation performance. Therefore, the present disclosure proposes a solution that can accelerate the loading speed of the computing unit and improve the operation performance.
[0028] Figure 2 A schematic diagram showing the processing operations of a processing device according to an embodiment of the present application is shown.
[0029] As Figure 2 shown, the first computing unit CU0 among the multiple computing units loads the first section a0 of the first tensor part A0 of the first tensor A into the first memory M0 corresponding to the first computing unit CU0, and copies the first section a0 of the first tensor part A0 to the second memory M1 corresponding to the second computing unit CU1 among the multiple computing units.
[0030] In this way, the first computing unit CU0 loads a part of the first tensor part A0, that is, the first section a0 (but not all), and through replication between computing units, the second computing unit CU1 can share the data of the first section a0 without the second computing unit CU1 loading the first section a0 from the outside. Since the replication speed between computing units is much faster than the external speed, and the amount of data loaded by the computing unit from the outside is less than the total amount of data originally to be loaded, a solution can greatly reduce the amount of data and time loaded by the computing unit, speed up the loading speed of the computing unit, and improve the operation performance.
[0031] The loading includes loading from the global memory, and the replication includes multicast. Of course, the loading can also include loading from other places, and the replication can also include other ways to spread data between computing units. Among them, the replication can be carried out simultaneously with the loading. For example, when the computing unit loads some data from the outside, the loaded data can be immediately copied to other computing units. Of course, a mechanism can also be designed to forward the data to other computing units while receiving the loaded data.
[0032] As Figure 2 shown, the second computing unit CU1 loads the second section a1 of the first tensor part A0 into the second memory M1 and copies the second section a1 of the first tensor part A0 into the first memory M0.
[0033] At this time, the memory M0 of the first computing unit CU0 has stored the combination of the first section a0 of the first tensor part A0 and the second section a1 of the first tensor part A0, that is, the first tensor part A0. The memory M1 of the second computing unit CU1 has stored the combination of the first section a0 of the first tensor part A0 and the second section a1 of the first tensor part A0, that is, the first tensor part A0.
[0034] That is to say, both the memory M0 of the first computing unit CU0 and the memory M1 of the second computing unit CU1 store the complete first tensor part A0.
[0035] In this way, through the loading and copying operations among different computing units, it is possible to greatly reduce the amount of data loaded by the computing units and the time, speed up the loading speed of the computing units, and improve the operation performance while ensuring that multiple computing units have the complete tensor parts they need for operation. Since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing units is overall accelerated and the operation performance is improved. Moreover, each computing unit can perform the loading and copying operations in parallel (without mutual dependence), so that each computing unit can complete the storage of the data to be operated in the memory faster, and thus can also start the operation on the stored data earlier and improve the operation performance.
[0036] Continue as Figure 2 shown, the first computing unit CU0 loads the first section b0 of the second tensor part B0 of the second tensor into the first memory M0, and copies the first section b0 of the second tensor part B0 to the third memory M2 corresponding to the third computing unit CU2 among the multiple computing units. The third computing unit CU2 loads the second section b1 of the second tensor part B0 into the third memory M2, and copies the second section b1 of the second tensor part B0 to the first memory M0.
[0037] At this time, the memory M0 of the first computing unit CU0 has stored the combination of the first section b0 of the second tensor part B0 and the second section b1 of the second tensor part B0, that is, the second tensor part B0. The third memory M2 of the third computing unit CU2 has stored the combination of the first section b0 of the second tensor part B0 and the second section b1 of the second tensor part B0, that is, the second tensor part B0.
[0038] That is to say, both the memory M0 of the first computing unit CU0 and the memory M2 of the third computing unit CU2 have stored the complete second tensor part B0.
[0039] Here, the first tensor part A0 and the second tensor part B0 that need to be operated have been stored in the memory M0 of the first computing unit CU0. Therefore, the first computing unit CU0 can perform operations on the first tensor part A0 and the second tensor part B0 stored in its memory M0, such as multiply-accumulate operations of matrix multiplication or convolution.
[0040] As Figure 2As shown, the third computing unit CU2 loads a first section a2 of another first tensor part A1 of the first tensor into the third memory M2, and copies the first section a2 of the another first tensor part A1 to a fourth memory M3 corresponding to a fourth computing unit CU3 among the multiple computing units. The fourth computing unit CU3 loads a second section a3 of the another first tensor part A1 into the fourth memory M3, and copies the second section a3 of the another first tensor part A1 to the third memory M2.
[0041] At this time, the memory M2 of the third computing unit CU2 has stored the combination of the first section a2 and the second section a3 of the another first tensor part A1, that is, the another first tensor part A1. The memory M3 of the fourth computing unit CU3 has stored the combination of the first section a2 and the second section a3 of the another first tensor part A1, that is, the another first tensor part A1.
[0042] That is to say, both the memory M2 of the third computing unit CU2 and the memory M3 of the fourth computing unit CU3 have stored the complete another first tensor part A1.
[0043] Here, the memory M2 of the third computing unit CU2 has stored the another first tensor part A1 and the second tensor part B0 that need to be operated on. Therefore, the third computing unit CU2 can perform operations on the another first tensor part A1 and the second tensor part B0 stored in its memory M2, such as multiply-accumulate operations for matrix multiplication or convolution.
[0044] As Figure 2 shown, the second computing unit CU1 loads a first section b2 of another second tensor part B1 of the second tensor into the second memory M1, and copies the first section b2 of the another second tensor part B1 to the fourth memory M3 corresponding to the fourth computing unit CU3. The fourth computing unit CU3 loads a second section b3 of the another second tensor part B1 into the fourth memory M3, and copies the second section b3 of the another second tensor part B1 to the second memory M1.
[0045] At this time, the memory M1 of the second computing unit CU1 has stored the combination of the first section b2 and the second section b3 of the another second tensor part B1, that is, the another second tensor part B1. The memory M3 of the fourth computing unit CU3 has stored the combination of the first section b2 and the second section b3 of the another second tensor part B1, that is, the another second tensor part B1.
[0046] That is to say, the memory M1 of the second computing unit CU1 and the memory M3 of the fourth computing unit CU3 both store a complete other second tensor part B1.
[0047] Here, the first tensor part A0 and the other second tensor part B1 that need to be operated on have already been stored in the memory M1 of the second computing unit CU1. Therefore, the second computing unit CU1 can perform operations on the first tensor part A0 and the other second tensor part B1 stored in its memory M1, such as multiply-accumulate operations for matrix multiplication or convolution.
[0048] Here, the other first tensor part A1 and the other second tensor part B1 that need to be operated on have already been stored in the memory M3 of the fourth computing unit CU3. Therefore, the fourth computing unit CU3 can perform operations on the other first tensor part A1 and the other second tensor part B1 stored in its memory M3, such as multiply-accumulate operations for matrix multiplication or convolution.
[0049] In this way, the first computing unit CU0, the second computing unit CU1, the third computing unit CU2, and the fourth computing unit CU3 can respectively perform operations on the data in the corresponding memories.
[0050] Note that the loading, copying, and operation operations of each computing unit can be performed in parallel respectively. For example, when the first computing unit performs its loading and copying operations, the second computing unit can simultaneously perform its loading and copying operations. Therefore, when the loading and copying operations of the first computing unit or the second computing unit are completed, both computing units have the corresponding data. Moreover, as long as the memory of the computing unit already has the data required for the operation, the computing unit can start the operation.
[0051] In this way, since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing units is overall accelerated and the operation performance is improved. Moreover, each computing unit can perform the loading and copying operations in parallel respectively (without depending on each other), so that each computing unit can more quickly complete the storage of the data to be operated in the memory, and thus can also start the operation on the stored data earlier and improve the operation performance.
[0052] Only 2*2 computing units are exemplified above, and the first tensor is divided into 2 first tensor parts, and each first tensor part is further divided into 2 segments, the second tensor is divided into 2 second tensor parts, and each second tensor part is further divided into 2 segments, so as to make full use of the 2*2 computing units, not only making full use of the loading and copying operations of the 2*2 computing units but also making full use of the subsequent operation operations of the 2*2 computing units.
[0053] However, the present disclosure is not limited to the above number of examples. The present disclosure also includes determining the size of each section and the splitting manner of the tensor according to the size of the memory corresponding to each computing unit and the number of computing units, so as to make full use of the resources of the computing units and the memory and improve the operation efficiency.
[0054] Suppose the first tensor is split into N first tensor parts, each first tensor part is split into M sections, and the second tensor is split into M second tensor parts, where each second tensor part is split into N sections, and N and M are positive integers; the plurality of computing units at least include an array of N rows * M columns of computing units. Note here that depending on different situations, N and M may not be equal or may be equal.
[0055] In one embodiment, the values of N and M are determined according to the size of the memory corresponding to each computing unit in the plurality of computing units and the number of the plurality of computing units, such that the size of each tensor part is not greater than the size of the memory divided by 2 and N * M is not greater than the number of the plurality of computing units, where if the sizes of the respective tensor parts are different, one or more tensor parts are supplemented with more than 1 zero to make the sizes of the respective tensor parts the same.
[0056] In this way, the size of each tensor part not being greater than the size of the memory divided by 2 can enable the memory of each computing unit to store the size of one first tensor part and the size of one second tensor part in total for performing operations between the two, and make full use of the memory space when the size of each tensor part is equal to the size of the memory divided by 2. If the sizes of the respective tensor parts are different, one or more tensor parts are supplemented with more than 1 zero to make the sizes of the respective tensor parts the same (if the size of the tensor part is exactly equal to the size of the memory divided by 2 and the sizes of the respective tensor parts are the same, then zero-padding is not required), to ensure that the sizes of the respective sections are equal and make full use of the memory space as much as possible.
[0057] At the same time, M computing units in one row can maximize the utilization of the resources of M computing units in one row by each loading and copying one section out of the M sections of the first tensor part to other computing units, so as to load data minimally and fastest and receive as much remaining data copied from other computing units as possible. N computing units in one column can maximize the utilization of the resources of N computing units in one column by each loading and copying one section out of the N sections of the second tensor part to other computing units, so as to load data minimally and fastest and receive as much remaining data copied from other computing units as possible.
[0058] In addition, the first tensor is divided into N first-tensor parts, and the second tensor is divided into M second-tensor parts. This enables each of the computing units in the array of all N rows * M columns of computing units to store one of the N first-tensor parts and one of the M second-tensor parts, and then perform operations on one of the N first-tensor parts and one of the M second-tensor parts in parallel.
[0059] That N*M is not greater than the number of the multiple computing units is also to ensure that there are sufficient computing units to perform the above operations.
[0060] Therefore, the determination and selection of N and M according to the size of the memory corresponding to each of the multiple computing units and the number of the multiple computing units can make full use of the resources of the computing units and the memory, and improve the operation efficiency.
[0061] After determining the size of the section and the values of N and M, the operations of loading, copying, and computing as described above can be performed.
[0062] Specifically, in one embodiment, the M computing units in the i-th row of the array of N rows * M columns of computing units respectively load one of the M sections in the i-th first-tensor part into the memory corresponding to themselves, and copy one of the M sections in the i-th first-tensor part to the memories corresponding to the remaining computing units other than themselves among the M computing units in the i-th row, so that the memories corresponding to the M computing units in the i-th row all have the i-th first-tensor part, where i is an integer in the range [1, N].
[0063] In one embodiment, the N computing units in the j-th column of the array of N rows * M columns of computing units respectively load one of the N sections in the j-th second-tensor part into the memory corresponding to themselves, and copy one of the N sections in the j-th second-tensor part to the memories corresponding to the remaining computing units other than themselves among the N computing units in the j-th column, so that the memories corresponding to the N computing units in the j-th column all have the j-th second-tensor part, where j is an integer in the range [1, M].
[0064] In one embodiment, the computing unit in the i-th row and j-th column of the array of N rows * M columns of computing units performs operations on the i-th first-tensor part and the j-th second-tensor part stored in the memory corresponding to itself.
[0065] In this way, each section of the corresponding tensor part in the first tensor is obtained by each row computing unit through external loading and copying from other computing units in the same row, and each section of the corresponding tensor part in the second tensor is obtained by each column computing unit through external loading and copying from other computing units in the same column. In this way, data can be loaded with the least amount and the fastest speed, and as much remaining data copied from other computing units can be received as possible, thereby accelerating the data loading speed of the computing units as a whole and improving the operation performance. Moreover, each computing unit can perform the operations of loading and copying separately in parallel (without mutual dependence), so that each computing unit can complete the storage of the data to be operated in the memory faster, and thus the operation on the stored data can also be started earlier and the operation performance can be improved.
[0066] Of course, although the above describes the correspondence between the rows and columns of the computing units and the first tensor and the second tensor in the form of an array of N rows * M columns of computing units, it is only for simplicity of description in this article and not a limitation. In reality, the computing units can also be arranged randomly (for example, in a row), but as long as these computing units can complete the above operations of loading, copying and computing.
[0067] Figure 3 FIG. shows a schematic diagram of the operation of an array of 2 rows * 4 columns (N = 2, M = 4) of computing units according to an embodiment of the present application.
[0068] For example, if the size of the memory corresponding to each computing unit is 128 kilobytes (KB), the size of the first tensor is 128 kilobytes (KB), and the size of the second tensor is 256 kilobytes (KB), then the number of computing units is 8. Therefore, the memory corresponding to each computing unit needs to load data from the first tensor and data from the second tensor, a total of 128 kilobytes (KB). Therefore, the data from the first tensor is at most 64 kilobytes (KB), and the data from the second tensor is at most 64 kilobytes (KB). Then the size of each part of the first tensor is at most 64 kilobytes (KB), and the size of each part of the second tensor is at most 64 kilobytes (KB). The size of the first tensor of 128 kilobytes (KB) can be divided into 2 parts of the first tensor, and the size of the second tensor of 256 kilobytes (KB) can be divided into 4 parts of the tensor. Therefore, it is determined that N = 2 and M = 4, and a total of 8 computing units are required.
[0069] Here, since the size of the first tensor and the size of the second tensor can be exactly divisible by 64, but if, for example, the size of the first tensor is 96 bits and cannot be divisible by 64, or in other words, it will result in two parts of the first tensor after splitting being 64 bits and 32 bits, then 32 zeros can be added to the 32-bit part of the first tensor, that is, the size of the first tensor part remains 64 bits. It's just that the first tensor part is 64 bits, and the second part of the first tensor consists of 32 bits and 32 zeros.
[0070] In this example, since N = 2 and M = 4, the first tensor is split into two first tensor parts A1 and A2, and each first tensor part is split into 4 segments: A1 is split into a1, a2, a3, a4, and A2 is split into a5, a6, a7, a8. The second tensor is split into 4 second tensor parts B1, B2, B3, B4, where each second tensor part is split into 2 segments: B1 is split into b1, b2, B2 is split into b3, b4, B3 is split into b5, b6, and B4 is split into b7, b8.
[0071] The 4 computing units in the i-th row of the array of 2 rows * 4 columns of computing units respectively load one of the 4 segments in the i-th first tensor part into the memory corresponding to themselves, and copy one of the 4 segments in the i-th first tensor part to the memories corresponding to the remaining computing units in the 4 computing units of the i-th row except themselves, so that the memories corresponding to the 4 computing units in the i-th row all have the i-th first tensor part, where i is an integer in the range [1, 2].
[0072] Specifically, as Figure 3 shown, the 4 computing units CU11, CU12, CU13, CU14 in the first row of the array of 2 rows * 4 columns of computing units respectively load one of the 4 segments a1, a2, a3, a4 in the first first tensor part into the memory corresponding to themselves, and copy one of the 4 segments in the first first tensor part to the memories corresponding to the remaining computing units in the 4 computing units of the first row except themselves, so that the memories corresponding to the 4 computing units in the first row all have the first first tensor part.
[0073] Specifically, the computing unit CU11 loads the section a1 into the corresponding memory M11 of its own, and copies the section a1 into the memories M12, M13, and M14 corresponding to the remaining computing units CU12, CU13, and CU14 respectively. The computing unit CU12 loads the section a2 into the corresponding memory M12 of its own, and copies the section a2 into the memories M11, M13, and M14 corresponding to the remaining computing units CU11, CU13, and CU14 respectively. The computing unit CU13 loads the section a3 into the corresponding memory M13 of its own, and copies the section a3 into the memories M11, M12, and M14 corresponding to the remaining computing units CU11, CU12, and CU14 respectively. The computing unit CU14 loads the section a4 into the corresponding memory M14 of its own, and copies the section a4 into the memories M11, M12, and M13 corresponding to the remaining computing units CU11, CU12, and CU13 respectively. So that the memories corresponding to the 4 computing units in the first row all have the sum A1 of the first tensor parts, that is, a1, a2, a3, and a4.
[0074] Each of the 4 computing units in the second row of the 2-row * 4-column array of computing units loads one of the 4 sections in the second first tensor part into the corresponding memory of its own, and copies one of the 4 sections in the second first tensor part into the memories corresponding to the remaining computing units other than itself among the 4 computing units in the second row, so that the memories corresponding to the 4 computing units in the second row all have the second first tensor part.
[0075] Specifically, the computing unit CU21 loads the section a5 into the corresponding memory M21 of its own, and copies the section a5 into the memories M22, M23, and M24 corresponding to the remaining computing units CU22, CU23, and CU24 respectively. The computing unit CU22 loads the section a6 into the corresponding memory M22 of its own, and copies the section a6 into the memories M21, M23, and M24 corresponding to the remaining computing units CU21, CU23, and CU24 respectively. The computing unit CU23 loads the section a7 into the corresponding memory M23 of its own, and copies the section a7 into the memories M21, M22, and M24 corresponding to the remaining computing units CU21, CU22, and CU24 respectively. The computing unit CU24 loads the section a8 into the corresponding memory M24 of its own, and copies the section a8 into the memories M21, M22, and M23 corresponding to the remaining computing units CU21, CU22, and CU23 respectively. So that the memories corresponding to the 4 computing units in the second row all have the sum A2 of the second first tensor parts, that is, a5, a6, a7, and a8.
[0076] In one embodiment, two computing units in the j-th column of the 2-row * 4-column array of computing units each load one of two segments in the j-th second tensor part into the memory corresponding to themselves, and copy one of the two segments in the j-th second tensor part to the memories corresponding to the remaining computing units among the four computing units in the j-th row, except themselves, so that the memories corresponding to the two computing units in the j-th column all have the j-th second tensor part, where j is an integer in the range [1, 4].
[0077] Specifically, as Figure 3 shown, two computing units CU11 and CU21 in the first column of the 2-row * 4-column array of computing units each load one of two segments b1 and b2 in the first second tensor part into the memory corresponding to themselves, and copy one of the two segments in the first second tensor part to the memories corresponding to the remaining computing units among the two computing units in the first column, except themselves, so that the memories corresponding to the two computing units in the first column all have the first second tensor part B1.
[0078] Specifically, computing unit CU11 loads segment b1 into the memory M11 corresponding to itself, and copies segment b1 to the memory M21 corresponding to the remaining computing unit CU21. Computing unit CU21 loads segment b2 into the memory M21 corresponding to itself, and copies segment b2 to the memory M11 corresponding to the remaining computing unit CU11. So that the memories corresponding to the two computing units in the first column all have the first second tensor part, that is, the sum B1 of b1 and b2.
[0079] Two computing units CU12 and CU22 in the second column of the 2-row * 4-column array of computing units each load one of two segments b3 and b4 in the second second tensor part into the memory corresponding to themselves, and copy one of the two segments in the second second tensor part to the memories corresponding to the remaining computing units among the two computing units in the second column, except themselves, so that the memories corresponding to the two computing units in the second column all have the second second tensor part B2.
[0080] Specifically, the computing unit CU12 loads the section b3 into the corresponding memory M12 of itself, and copies the section b3 to the memory M22 corresponding to the remaining computing units CU22. The computing unit CU22 loads the section b4 into the corresponding memory M22 of itself, and copies the section b4 to the memory M12 corresponding to the remaining computing units CU12. So that the memories corresponding to the two computing units in the second column both have the sum B2 of the second tensor parts b3 and b4, i.e., the second second tensor part.
[0081] Two computing units CU13 and CU23 in the third column of the array of the 2 rows * 4 columns of computing units respectively load one of the two sections b5 and b6 in the third second tensor part into the corresponding memory of themselves, and copy one of the two sections in the third second tensor part to the memories corresponding to the remaining computing units other than themselves among the two computing units in the second column, so that the memories corresponding to the two computing units in the third column both have the third second tensor part B3.
[0082] Specifically, the computing unit CU13 loads the section b5 into the corresponding memory M13 of itself, and copies the section b5 to the memory M23 corresponding to the remaining computing units CU23. The computing unit CU23 loads the section b6 into the corresponding memory M23 of itself, and copies the section b6 to the memory M13 corresponding to the remaining computing units CU13. So that the memories corresponding to the two computing units in the third column both have the sum B3 of the third second tensor parts b5 and b6, i.e., the third second tensor part.
[0083] Two computing units CU14 and CU24 in the fourth column of the array of the 2 rows * 4 columns of computing units respectively load one of the two sections b7 and b8 in the fourth second tensor part into the corresponding memory of themselves, and copy one of the two sections in the fourth second tensor part to the memories corresponding to the remaining computing units other than themselves among the two computing units in the second column, so that the memories corresponding to the two computing units in the fourth column both have the fourth second tensor part B4.
[0084] Specifically, the computing unit CU14 loads the section b7 into the corresponding memory M14 of itself, and copies the section b7 to the memory M24 corresponding to the remaining computing units CU24. The computing unit CU24 loads the section b8 into the corresponding memory M24 of itself, and copies the section b8 to the memory M14 corresponding to the remaining computing units CU14. So that the memories corresponding to the two computing units in the fourth column both have the sum B4 of the fourth second tensor parts b7 and b8, i.e., the fourth second tensor part.
[0085] Therefore, the computing unit at the i-th row and j-th column in the array of 2 rows * 4 columns of computing units operates on the i-th first tensor part stored in the corresponding memory of itself and the j-th second tensor part. For example, the computing unit CU11 operates on the first first tensor part A1 and the first second tensor part B1, the computing unit CU12 operates on the first first tensor part A1 and the second second tensor part B2, the computing unit CU13 operates on the first first tensor part A1 and the third second tensor part B3, and the computing unit CU14 operates on the first first tensor part A1 and the fourth second tensor part B4. The computing unit CU21 operates on the second first tensor part A2 and the first second tensor part B1, the computing unit CU22 operates on the second first tensor part A2 and the second second tensor part B2, the computing unit CU23 operates on the second first tensor part A2 and the third second tensor part B3, and the computing unit CU24 operates on the second first tensor part A2 and the fourth second tensor part B4.
[0086] In this way, each section of the corresponding tensor part in the first tensor is obtained by each row of computing units through external loading and copying from other computing units in the same row, and each section of the corresponding tensor part in the second tensor is obtained by each column of computing units through external loading and copying from other computing units in the same column. In this way, data can be loaded with the least amount and the fastest speed, and as much remaining data copied from other computing units as possible can be received, thereby accelerating the data loading speed of the computing units as a whole and improving the operation performance. Moreover, each computing unit can perform the operations of loading and copying separately in parallel (without depending on each other), so that each computing unit can complete the storage of the data to be operated in the memory faster, and thus the operation on the stored data can also be started earlier and the operation performance can be improved.
[0087] The present disclosure also provides a synchronization mechanism for a memory to receive data and consume data.
[0088] Specifically, in one embodiment, each of the multiple computing units records a first count related to the size of the data received in the memory corresponding to each computing unit. Wherein, in response to the first count satisfying a predetermined condition, the computing unit starts to operate on the received data, and the first count is cleared.
[0089] In one embodiment, the predetermined condition includes at least one of the following: the first count is equal to the size of the memory corresponding to each of the multiple computing units; or the first count is equal to the size of the data to be processed by each computing unit. Of course, the predetermined condition can also be other conditions. For example, a predetermined threshold is set in advance, and when the first count is equal to or greater than the predetermined threshold, it is considered that the predetermined condition is satisfied. The predetermined threshold can be less than or equal to the size of the memory.
[0090] In one embodiment, the memory fence of each computing unit is used to count the data received in the memory corresponding to each computing unit. The memory fence increments the first count by 1 each time the memory receives (including loading from the outside and obtaining from other computing units) one bit. In response to the first count satisfying the predetermined condition, the computing unit starts to process the received data. At this time, the first count can be cleared.
[0091] In this way, it is possible to determine whether the data in the memory is ready for processing, and when it is determined that the data in the memory is ready for processing, start to process the received data.
[0092] In one embodiment, each of the multiple computing units records a second count of the size of the data related to the first tensor read from the memory corresponding to each computing unit and a third count of the size of the data related to the second tensor read from the memory corresponding to each computing unit. Wherein, in response to the second count and the third count satisfying the predetermined condition, the computing unit starts to perform other recording, copying, and processing operations, and the second count is cleared, and the third count is cleared.
[0093] In one embodiment, the predetermined condition includes at least one of the following: the sum of the second count and the third count is equal to the size of the memory corresponding to each of the multiple computing units; or the sum of the second count and the third count is equal to the size of the data for which each computing unit performs processing.
[0094] In one embodiment, each computing unit increments the second count by 1 when receiving one bit related to the first tensor, and each computing unit increments the third count by 1 when receiving one bit related to the second tensor.
[0095] Then, in response to the second count and the third count satisfying the predetermined condition, the computing unit starts to perform other recording, copying, and processing operations, and the second count is cleared, and the third count is cleared. At this time, the second count and the third count can be cleared.
[0096] In this way, it is possible to determine whether the memory is idle and ready to receive a new round of data, and in the case where it is determined that the memory is idle and ready to receive a new round of data, other recording, copying, and computing operations are started to reuse the memory again.
[0097] In this way, the synchronization problem of the memory receiving data and the memory consuming data during the above loading, copying, and computing processes can be solved.
[0098] Figure 4A A flowchart of a processing method for a processing device according to an embodiment of the present application is shown.
[0099] The processing device includes a plurality of computing units and a plurality of memories corresponding to the plurality of computing units. As Figure 4A shown, the processing method for the processing device includes: Step 410, the first computing unit CU0 among the plurality of computing units loads the first segment a0 of the first tensor part A0 of the first tensor into the first memory M0 corresponding to the first computing unit CU0; Step 420, the first computing unit CU0 copies the first segment a0 of the first tensor part A0 to the second memory M1 corresponding to the second computing unit CU1 among the plurality of computing units.
[0100] In this way, the first computing unit CU0 performs the loading of only a part of the first tensor part A0, that is, the first segment a0 instead of all. Through the copying between the computing units, the second computing unit CU1 can share the data of the first segment a0 without the second computing unit CU1 loading the first segment a0 from the outside. And since the speed of copying between the computing units is much faster than the speed from the outside, and the amount of data loaded by the computing unit from the outside is less than the total amount of data to be originally loaded, therefore, a solution can be achieved that can greatly reduce the amount of data and time loaded by the computing unit, speed up the loading speed of the computing unit, and improve the computing performance.
[0101] Figure 4B A flowchart of a processing method for a processing device according to another embodiment of the present application is shown.
[0102] As Figure 4A shown, the processing method for the processing device includes: Step 410, the first computing unit CU0 among the plurality of computing units loads the first segment a0 of the first tensor part A0 of the first tensor into the first memory M0 corresponding to the first computing unit CU0; Step 420, the first computing unit CU0 copies the first segment a0 of the first tensor part A0 to the second memory M1 corresponding to the second computing unit CU1 among the plurality of computing units.
[0103] In one embodiment, the processing method further includes: Step 430, the second computing unit CU1 loads the second section a1 of the first tensor part A0 into the second memory M1, and copies the second section a1 of the first tensor part A0 into the first memory M0.
[0104] In this way, through the mutual loading and copying operations of different computing units, it is possible to greatly reduce the data volume and time loaded by the computing units, speed up the loading speed of the computing units, and improve the operation performance while ensuring that multiple computing units have the complete tensor parts they need for operation. Since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing units is overall accelerated and the operation performance is improved. Moreover, each computing unit can perform the loading and copying operations in parallel respectively without depending on each other. Therefore, it is possible to make each computing unit complete the storage of the data to be operated in the memory faster, and thus it is also possible to start the operation on the stored data earlier and improve the operation performance.
[0105] Figure 4C The flowchart of a processing method for a processing device according to another embodiment of the present application is shown.
[0106] As Figure 4C shown, in addition to Figure 4B the steps 410-430 shown, the processing method further includes: Step 440, the first computing unit CU0 loads the first section b0 of the second tensor part B0 of the second tensor into the first memory M0, and copies the first section b0 of the second tensor part B0 into the third memory M2 corresponding to the third computing unit CU2 among the multiple computing units. The third computing unit CU2 loads the second section b1 of the second tensor part B0 into the third memory M2, and copies the second section b1 of the second tensor part B0 into the first memory M0.
[0107] The processing method further includes: Step 450, the third computing unit CU2 loads the first section a2 of another first tensor part A1 of the first tensor into the third memory M2, and copies the first section a2 of the another first tensor part A1 into the fourth memory M3 corresponding to the fourth computing unit CU3 among the multiple computing units. The fourth computing unit CU3 loads the second section a3 of the another first tensor part A1 into the fourth memory M3, and copies the second section a3 of the another first tensor part A1 into the third memory M2.
[0108] The processing method further includes: Step 460, the second computing unit CU1 loads a first section b2 of another second tensor part B1 of the second tensor into the second memory M1, and copies the first section b2 of the another second tensor part B1 to the fourth memory M3 corresponding to the fourth computing unit CU3. The fourth computing unit CU3 loads a second section b3 of the another second tensor part B1 into the fourth memory M3, and copies the second section b3 of the another second tensor part B1 to the second memory M1.
[0109] The processing method further includes: Step 470, the first computing unit CU0, the second computing unit CU1, the third computing unit CU2, and the fourth computing unit CU3 respectively perform operations on the data in the corresponding memories.
[0110] In this way, since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing units is overall accelerated and the operation performance is improved. Moreover, each computing unit can perform the operations of loading and copying in parallel respectively without depending on each other. Therefore, it is possible to more quickly store the data to be operated by each computing unit in the memory, and thus it is also possible to start the operation on the stored data earlier and improve the operation performance.
[0111] In one embodiment, the first tensor is divided into N first tensor parts, each first tensor part is divided into M sections, the second tensor is divided into M second tensor parts, where each second tensor part is divided into N sections, where N and M are positive integers; the plurality of computing units at least include an array of N rows * M columns of computing units.
[0112] Figure 5 The flowchart of a processing method for a processing device according to another embodiment of the present application is shown.
[0113] The processing method includes:
[0114] Step 510, M computing units in the i-th row of the array of N rows * M columns of computing units respectively load one section of the M sections in the i-th first tensor part into the memory corresponding to themselves, and copy one section of the M sections in the i-th first tensor part to the memories respectively corresponding to the remaining computing units except themselves among the M computing units in the i-th row, so that the memories corresponding to the M computing units in the i-th row all have the i-th first tensor part, where i is an integer in the range [1, N].
[0115] Step 520: The N computing units in the j-th column of the array of N rows * M columns of computing units respectively load one of the N segments in the j-th second tensor part into the memory corresponding to themselves, and copy one of the N segments in the j-th second tensor part to the memories corresponding to the remaining computing units in the j-th column of the N computing units except themselves, so that the memories corresponding to the N computing units in the j-th column all have the j-th second tensor part, where j is an integer in the range [1, M].
[0116] The processing method further includes: Step 530: The computing unit in the i-th row and j-th column of the array of N rows * M columns of computing units performs an operation on the i-th first tensor part stored in the memory corresponding to itself and the j-th second tensor part.
[0117] In this way, each segment of the corresponding tensor part in the first tensor is obtained by each row of computing units through external loading and copying from other computing units in the same row, and each segment of the corresponding tensor part in the second tensor is obtained by each column of computing units through external loading and copying from other computing units in the same column. In this way, data can be loaded with the least amount and the fastest speed, and as much remaining data copied from other computing units as possible can be received, thereby overall accelerating the data loading speed of the computing units and improving the operation performance. Moreover, each computing unit can perform the operations of loading and copying separately and in parallel without mutual dependence. Therefore, it can be made faster for each computing unit to complete the storage of the data to be operated in the memory, and thus the operation on the stored data can also be started earlier and the operation performance can be improved.
[0118] In one embodiment, the values of N and M are determined according to the size of the memory corresponding to each of the multiple computing units and the number of the multiple computing units, such that the size of each tensor part is not greater than the size of the memory divided by 2 and N * M is not greater than the number of the multiple computing units. Where if the sizes of the respective tensor parts are different, one or more tensor parts are supplemented with more than 1 zero to make the sizes of the respective tensor parts the same.
[0119] Thus, the size of each tensor part not being greater than the size of the memory divided by 2 enables the memory of each computing unit to store the size of a first tensor part and the size of a second tensor part in total for operations between the two, and the space of the memory is fully utilized when the size of each tensor part is equal to the size of the memory divided by 2. If the sizes of the respective tensor parts are different, one or more zeros are added to one or more tensor parts to make the sizes of the respective tensor parts the same. If the size of the tensor part is exactly equal to the size of the memory divided by 2 and the sizes of the respective tensor parts are the same, no zero-padding is required to ensure that the sizes of the respective sections are equal and the space of the memory is utilized as fully as possible.
[0120] Figure 6 The flowchart shows the synchronization mechanism of the memory receiving data and the memory consuming data included in the processing method according to an embodiment of the present application.
[0121] The processing method further includes: Step 610, each of the multiple computing units records a first count related to the size of the data received in the memory corresponding to each computing unit, wherein, in response to the first count satisfying a predetermined condition, the computing unit starts to operate on the received data, and the first count is cleared.
[0122] In one embodiment, the predetermined condition includes at least one of the following: the first count is equal to the size of the memory corresponding to each of the multiple computing units; or the first count is equal to the size of the data to be operated on by each computing unit.
[0123] In one embodiment, the memory fence of each computing unit is used to count the data received in the memory corresponding to each computing unit, wherein the memory fence increments the first count by 1 each time one bit is received by the memory.
[0124] Thus, it is possible to determine whether the data in the memory is ready for operation, and when it is determined that the data in the memory is ready for operation, the operation on the received data is started.
[0125] The processing method further includes: Step 620, each of the multiple computing units records a second count of the size of the data related to the first tensor read from the memory corresponding to each computing unit and a third count of the size of the data related to the second tensor read from the memory corresponding to each computing unit, wherein, in response to the second count and the third count satisfying a predetermined condition, the computing unit starts other operations of recording, copying, and operating, and the second count is cleared and the third count is cleared.
[0126] In one embodiment, the predetermined condition includes at least one of the following: the sum of the second count and the third count is equal to the size of the memory corresponding to each of the plurality of computing units; or the sum of the second count and the third count is equal to the size of the data for which each computing unit performs operations.
[0127] In one embodiment, each computing unit increments the second count by 1 when receiving one bit related to the first tensor, and each computing unit increments the third count by 1 when receiving one bit related to the second tensor.
[0128] In this way, it is possible to determine whether the memory is idle and ready to receive a new round of data, and in the case where it is determined that the memory is idle and ready to receive a new round of data, other operations such as recording, copying, and computing are started to reuse the memory again.
[0129] In this way, the synchronization problem of the memory receiving data and the memory consuming data in the above loading, copying, and computing processes can be solved.
[0130] In one embodiment, the operation includes matrix multiplication operation, where the first tensor is the left matrix and the second tensor is the right matrix.
[0131] In one embodiment, the operation includes convolution operation, where the first tensor is the input tensor and the second tensor is the weight tensor.
[0132] In one embodiment, loading includes loading from a global memory, and the copying includes multicasting.
[0133] Figure 7 A block diagram of an electronic device according to an embodiment of the present application is shown.
[0134] As Figure 7 shown, the electronic device 700 includes: a loading device 710 configured to load a first section a0 of a first tensor portion A0 of the first tensor into a first memory M0 corresponding to the first computing unit CU0 by the first computing unit CU0 among the plurality of computing units; a copying device 720 configured to copy the first section a0 of the first tensor portion A0 to a second memory M1 corresponding to a second computing unit CU1 among the plurality of computing units by the first computing unit CU0.
[0135] In this way, the first computing unit CU0 loads a part of the first tensor part A0, that is, the first section a0, instead of the whole, and through the copying between computing units, the second computing unit CU1 can share the data of the first section a0 without loading the first section a0 from the outside. Since the speed of copying between computing units is much faster than that from the outside, and the amount of data loaded by the computing unit from the outside is less than the total amount of data to be loaded originally, a solution can greatly reduce the amount of data and time loaded by the computing unit, speed up the loading speed of the computing unit and improve the operation performance.
[0136] The electronic device 700 further includes: a first device (not shown in the figure), configured to load, by the second computing unit CU1, a second section a1 of the first tensor part A0 into the second memory M1, and copy the second section a1 of the first tensor part A0 into the first memory M0.
[0137] In this way, through the loading and copying operations between different computing units, a solution can greatly reduce the amount of data and time loaded by the computing unit, speed up the loading speed of the computing unit and improve the operation performance while ensuring that multiple computing units have the complete tensor parts required for their operations. Since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing unit can be overall speeded up and the operation performance can be improved. Moreover, each computing unit can perform the loading and copying operations in parallel respectively without depending on each other, so that each computing unit can complete the storage of the data to be operated in the memory faster, and thus can also start the operation on the stored data earlier and improve the operation performance.
[0138] Figure 4C The flowchart of a processing method for a processing device according to another embodiment of the present application is shown.
[0139] The electronic device 700 further includes: a second device (not shown in the figure), configured to load, by the first computing unit CU0, a first section b0 of the second tensor part B0 of the second tensor into the first memory M0, and copy the first section b0 of the second tensor part B0 into the third memory M2 corresponding to the third computing unit CU2 among the multiple computing units, and load, by the third computing unit CU2, a second section b1 of the second tensor part B0 into the third memory M2, and copy the second section b1 of the second tensor part B0 into the first memory M0.
[0140] The electronic device 700 further includes: a third device (not shown in the figure), configured to load, by the third computing unit CU2, a first section a2 of another first tensor part A1 of the first tensor into the third memory M2, and copy the first section a2 of the another first tensor part A1 to a fourth memory M3 corresponding to a fourth computing unit CU3 among the plurality of computing units, and load, by the fourth computing unit CU3, a second section a3 of the another first tensor part A1 into the fourth memory M3, and copy the second section a3 of the another first tensor part A1 to the third memory M2.
[0141] The electronic device 700 further includes: a fourth device (not shown in the figure), configured to load, by the second computing unit CU1, a first section b2 of another second tensor part B1 of the second tensor into the second memory M1, and copy the first section b2 of the another second tensor part B1 to a fourth memory M3 corresponding to the fourth computing unit CU3, and load, by the fourth computing unit CU3, a second section b3 of the another second tensor part B1 into the fourth memory M3, and copy the second section b3 of the another second tensor part B1 to the second memory M1.
[0142] The electronic device 700 further includes: a fifth device (not shown in the figure), configured to perform operations on the data in the corresponding memories by the first computing unit CU0, the second computing unit CU1, the third computing unit CU2, and the fourth computing unit CU3 respectively.
[0143] In this way, since each computing unit loads a part of the data and copies a part of the data from other computing units, the data loading speed of the computing units is overall accelerated and the operation performance is improved. Moreover, each computing unit can perform the operations of loading and copying in parallel respectively without depending on each other. Therefore, it is possible to more quickly enable each computing unit to complete the storage of the data to be operated in the memory, and thus it is also possible to start the operation on the stored data earlier and improve the operation performance.
[0144] In one embodiment, the first tensor is sliced into N first tensor parts, each first tensor part is sliced into M sections, the second tensor is sliced into M second tensor parts, where each second tensor part is sliced into N sections, where N and M are positive integers; the plurality of computing units includes at least an array of N rows * M columns of computing units.
[0145] The electronic device 700 further includes: a sixth device (not shown in the figure), configured to load, by each of the M computing units in the i-th row of the array of N rows * M columns of computing units, one of the M segments in the i-th first tensor part into the memory corresponding to itself, and copy one of the M segments in the i-th first tensor part to the memories respectively corresponding to the remaining computing units other than itself among the M computing units in the i-th row, so that the memories corresponding to the M computing units in the i-th row all have the i-th first tensor part, where i is an integer in the range [1, N].
[0146] The electronic device 700 further includes: a seventh device (not shown in the figure), configured to load, by each of the N computing units in the j-th column of the array of N rows * M columns of computing units, one of the N segments in the j-th second tensor part into the memory corresponding to itself, and copy one of the N segments in the j-th second tensor part to the memories respectively corresponding to the remaining computing units other than itself among the N computing units in the j-th column, so that the memories corresponding to the N computing units in the j-th column all have the j-th second tensor part, where j is an integer in the range [1, M].
[0147] The electronic device 700 further includes: an eighth device (not shown in the figure), configured to perform an operation on the i-th first tensor part and the j-th second tensor part stored in the memory corresponding to itself by the computing unit at the i-th row and j-th column of the array of N rows * M columns of computing units.
[0148] In this way, each row of computing units obtains the corresponding tensor part in the first tensor by loading from the outside and copying from other computing units in the same row for each segment of the corresponding tensor part in the first tensor, and each column of computing units obtains the corresponding tensor part in the second tensor by loading from the outside and copying from other computing units in the same column for each segment of the corresponding tensor part in the second tensor. In this way, data can be loaded with the least amount and the fastest speed, and as much remaining data copied from other computing units as possible can be received, thereby overall accelerating the data loading speed of the computing units and improving the operation performance. Moreover, each computing unit can perform the operations of loading and copying separately and in parallel without depending on each other, so that each computing unit can complete the storage of the data to be operated in the memory faster, and thus the operation on the stored data can also be started earlier and the operation performance can be improved.
[0149] In one embodiment, the values of N and M are determined according to the size of the memory corresponding to each of the multiple computing units and the number of the multiple computing units, such that the size of each tensor part is not greater than the size of the memory divided by 2 and N * M is not greater than the number of the multiple computing units, wherein if the sizes of the respective tensor parts are different, one or more zeros are added to one or more tensor parts to make the sizes of the respective tensor parts the same.
[0150] In this way, the size of each tensor part not being greater than the size of the memory divided by 2 enables the memory of each computing unit to store the size of a first tensor part and the size of a second tensor part in total for performing an operation between the two, and the space of the memory is fully utilized when the size of each tensor part is equal to the size of the memory divided by 2. If the sizes of the respective tensor parts are different, one or more zeros are added to one or more tensor parts to make the sizes of the respective tensor parts the same. If the size of the tensor part is exactly equal to the size of the memory divided by 2 and the sizes of the respective tensor parts are the same, then no zero needs to be added to ensure that the sizes of all sections are equal and the space of the memory is utilized as fully as possible.
[0151] The electronic device 700 further includes: a ninth device (not shown in the figure), configured to record, by each of the multiple computing units, a first count related to the size of the data received in the memory corresponding to each computing unit, wherein in response to the first count satisfying a predetermined condition, the computing unit starts to perform an operation on the received data, and the first count is cleared.
[0152] In one embodiment, the predetermined condition includes at least one of the following: the first count is equal to the size of the memory corresponding to each of the multiple computing units; or the first count is equal to the size of the data to be operated on by each computing unit.
[0153] In one embodiment, the data received in the memory corresponding to each computing unit is counted through a memory fence of each computing unit, wherein the memory fence increments the first count by 1 each time one bit is received by the memory.
[0154] In this way, it is possible to determine whether the data in the memory is ready for operation, and when it is determined that the data in the memory is ready for operation, the operation on the received data is started.
[0155] The electronic device 700 further includes: a tenth device (not shown in the figure), configured to record, by each of the multiple computing units, a second count of the data size read from the memory corresponding to each computing unit related to the first tensor and a third count of the data size read from the memory corresponding to each computing unit related to the second tensor. Wherein, in response to the second count and the third count satisfying a predetermined condition, the computing unit starts operations of other recording, copying, and computing, and the second count is cleared, and the third count is cleared.
[0156] In one embodiment, the predetermined condition includes at least one of the following: the sum of the second count and the third count is equal to the size of the memory corresponding to each of the multiple computing units; or the sum of the second count and the third count is equal to the size of the data for each computing unit to perform operations.
[0157] In one embodiment, each computing unit increments the second count by 1 when receiving one bit related to the first tensor, and each computing unit increments the third count by 1 when receiving one bit related to the second tensor.
[0158] In this way, it is possible to determine whether the memory is idle and ready to receive a new round of data, and in the case where it is determined that the memory is idle and ready to receive a new round of data, start operations of other recording, copying, and computing to reuse the memory again.
[0159] In this way, it is possible to solve the synchronization problem of the memory receiving data and the memory consuming data during the above-mentioned loading, copying, and computing processes.
[0160] In one embodiment, the operation includes matrix multiplication operation, where the first tensor is the left matrix and the second tensor is the right matrix.
[0161] In one embodiment, the operation includes convolution operation, where the first tensor is the input tensor and the second tensor is the weight tensor.
[0162] In one embodiment, the loading includes loading from the global memory, and the copying includes multicast.
[0163] Figure 8 A block diagram of an exemplary electronic device suitable for implementing the embodiments of the present application is shown.
[0164] The electronic device may include a processor 810 and a storage medium 820. The storage medium 820 is coupled to the processor 810 and stores computer-executable instructions therein for performing the steps of the various methods of the embodiments of the present application when executed by the processor.
[0165] The processor 810 may include, but is not limited to, for example, one or more processors or microprocessors, etc.
[0166] The storage medium 820 may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).
[0167] In addition, the electronic device may further include (but is not limited to) a data bus 830, an input / output (I / O) bus 840, a display 805, and input / output devices 860 (such as a keyboard, a mouse, a speaker, etc.), etc.
[0168] The processor 810 may communicate with an external display 850 and input / output devices 860, etc. via the I / O bus 840 through a wired or wireless network (not shown).
[0169] The storage medium 820 may also store at least one computer-executable instruction for performing the steps of the various functions and / or methods in the embodiments described in the present technology when executed by the processor 810.
[0170] In one embodiment, the at least one computer-executable instruction may also be compiled into or constitute a computer program product or software product, and when one or more computer-executable instructions are executed by a processor, the steps of the various functions and / or methods in the embodiments described in the present technology are executed.
[0171] Figure 9 A schematic diagram of a non-transitory computer-readable storage medium according to an embodiment of the present application is shown.
[0172] As Figure 9 shown, instructions, such as computer instructions 910, are stored on the non-transitory computer-readable storage medium 920. When the computer instructions 910 are executed by a processor, the various methods described above may be executed. The non-transitory computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. Non-transitory non-volatile memory may include, for example, read-only memory (ROM), hard disks, flash memory, etc. For example, the non-transitory computer-readable storage medium 920 may be connected to a computing device such as a computer, and then, when the computing device executes the computer instructions 910 stored on the non-transitory computer-readable storage medium 920, the various methods described above may be performed.
[0173] The block diagrams of the devices, apparatuses, equipment, and systems involved in this disclosure are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including", "comprising", "having", etc. are open-ended terms, meaning "including but not limited to", and can be used interchangeably with each other. The phrase "such as / for example" used herein means "such as / for example but not limited to", and can be used interchangeably with it.
[0174] The flowchart of steps in this disclosure and the above method descriptions are only illustrative examples and are not intended to require or imply that the steps of each embodiment must be performed in the order given. As those skilled in the art will recognize, the steps in the above embodiments can be performed in any order. Words such as "subsequently", "then", "next", etc. are not intended to limit the order of the steps; these words are only used to guide the reader through the description of these methods. In addition, any reference to a singular element using articles "a", "an", or "the" is not to be construed as limiting that element to the singular.
[0175] In addition, the steps and apparatuses in each of the embodiments herein are not limited to being implemented in a particular embodiment. In fact, according to the concepts of this application, relevant partial steps and partial apparatuses in each of the embodiments herein can be combined to conceive new embodiments, and these new embodiments are also within the scope of this application.
[0176] The methods disclosed herein include steps for implementing the described methods. The methods and / or steps can be interchanged with each other without departing from the scope of the claims. In other words, unless a specific order of steps is specified, the order and / or use of the specific steps can be modified without departing from the scope of the claims.
[0177] The above methods can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the method can be stored as instructions on a tangible computer-readable medium. The storage medium can be any available tangible medium accessible by a computer. By way of example and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM, or other optical disc storage, magnetic disk storage, or other magnetic storage devices, or any other tangible medium that can be used to carry or store the desired program code in the form of instructions or data structures and that can be accessed by a computer. As used herein, disk and disc include compact disc (CD), laser disc, optical disc, digital versatile disc (DVD), floppy disk, and Blu-ray disc, where disk generally reproduces data magnetically, while disc uses lasers to optically reproduce data.
[0178] Accordingly, the present disclosure may also include a computer program product, wherein the computer program product can perform the methods, steps, and operations given herein. For example, such a computer program product may be a computer software package, computer code instructions, a computer-readable tangible medium having computer instructions tangibly stored (and / or encoded) thereon, which instructions are executable by a processor to perform the operations described herein. The computer program product may include packaging material.
[0179] In addition, modules and / or other suitable means for performing the methods and techniques described herein may be downloaded wirelessly from a server, when appropriate. Alternatively, the various methods described herein may be provided via a storage component so as to obtain the various methods when coupled to the storage component. Further, any other suitable techniques for providing the methods and techniques described herein to a device may be utilized.
[0180] The foregoing description has been presented for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the present application to the form disclosed herein. Although several example aspects and embodiments have been discussed above, those skilled in the art will recognize some of their variations, modifications, alterations, additions, and subcombinations.
Claims
1. A processing device comprising: A plurality of computing units and a plurality of memories corresponding to the plurality of computing units; The first tensor is divided into N first tensor parts, each of which is divided into M segments, the second tensor is divided into M second tensor parts, each of which is divided into N segments, where N and M are positive integers; the plurality of computing units at least includes an array of N rows and M columns of computing units, The M computing units in the i-th row of the array of N rows*M columns of computing units respectively load one of the M segments in the i-th first tensor part into the memory corresponding to themselves, and copy one of the M segments in the i-th first tensor part to the memory corresponding to each of the remaining computing units in the i-th row of M computing units except themselves, so that the memories corresponding to the M computing units in the i-th row all have the i-th first tensor part, where i is an integer in the range [1, N], The N computing units in the j-th column in the array of N rows and M columns of computing units respectively load one segment of the N segments in the j-th second tensor part into the memory corresponding to themselves, and copy one segment of the N segments in the j-th second tensor part to the memories corresponding to the remaining computing units among the N computing units in the j-th column except themselves, so that the memories corresponding to the N computing units in the j-th column all have the j-th second tensor part, where j is an integer in the range [1, M].
2. The processing device according to claim 1, wherein: A first computing unit (CU0) among the plurality of computing units loads a first segment (a0) of a first tensor portion (A0) of a first tensor into a first memory (M0) corresponding to the first computing unit (CU0), and copies the first segment (a0) of the first tensor portion (A0) to a second memory (M1) corresponding to a second computing unit (CU1) among the plurality of computing units.
3. The processing device according to claim 2, wherein the second computing unit (CU1) loads the second segment (a1) of the first tensor part (A0) into the second memory (M1) and copies the second segment (a1) of the first tensor part (A0) into the first memory (M0).
4. The processing device according to claim 3, wherein the first computing unit (CU0) loads a first segment (b0) of a second tensor part (B0) of a second tensor into the first memory (M0), and copies the first segment (b0) of the second tensor part (B0) to a third memory (M2) corresponding to a third computing unit (CU2) among the plurality of computing units, and the third computing unit (CU2) loads a second segment (b1) of the second tensor part (B0) into the third memory (M2), and copies the second segment (b1) of the second tensor part (B0) to the first memory (M0).
5. The processing device according to claim 4, wherein: The third computing unit (CU2) loads a first segment (a2) of another first tensor part (A1) of the first tensor into the third memory (M2), and copies the first segment (a2) of the another first tensor part (A1) to a fourth memory (M3) corresponding to a fourth computing unit (CU3) among the plurality of computing units, and the fourth computing unit (CU3) loads a second segment (a3) of the another first tensor part (A1) into the fourth memory (M3), and copies the second segment (a3) of the another first tensor part (A1) to the third memory (M2), The second computing unit (CU1) loads a first segment (b2) of another second tensor part (B1) of the second tensor into the second memory (M1), and copies the first segment (b2) of the another second tensor part (B1) to a fourth memory (M3) corresponding to the fourth computing unit (CU3); the fourth computing unit (CU3) loads a second segment (b3) of the another second tensor part (B1) into the fourth memory (M3), and copies the second segment (b3) of the another second tensor part (B1) to the second memory (M1).
6. The processing device according to claim 5, wherein: The first computing unit (CU0), the second computing unit (CU1), the third computing unit (CU2), and the fourth computing unit (CU3) respectively perform operations on data in corresponding memories.
7. The processing device according to claim 1, wherein: The computing unit in the i-th row and j-th column in the array of N rows*M columns of computing units operates on the i-th first tensor part and the j-th second tensor part stored in its corresponding memory.
8. The processing device according to claim 1, wherein: The values of N and M are determined according to the size of the memory corresponding to each of the multiple computing units and the number of the multiple computing units, so that the size of each tensor part is not greater than the size of the memory divided by 2 and N*M is not greater than the number of the multiple computing units, wherein if the sizes of the various tensor parts are different, one or more tensor parts are supplemented with one or more zeros to make the sizes of the various tensor parts the same.
9. The processing device according to claim 1, wherein: Each of the multiple computing units records a first count related to the size of data received in the memory corresponding to each computing unit, wherein, in response to the first count satisfying a predetermined condition, the computing unit starts to operate on the received data and the first count is cleared.
10. The processing device according to claim 9, wherein: The predetermined condition includes at least one of the following: the first count is equal to the size of the memory corresponding to each of the multiple computing units; or the first count is equal to the size of data to be calculated by each computing unit.
11. The processing device according to claim 9, wherein: Data received in a memory corresponding to each computing unit is counted by a memory fence of each computing unit, wherein the memory fence increases a first count by 1 each time a bit is received in the memory.
12. The processing device according to claim 1, wherein: Each of the multiple computing units records a second count of the size of data related to the first tensor read from the memory corresponding to each computing unit and a third count of the size of data related to the second tensor read from the memory corresponding to each computing unit, wherein, in response to the second count and the third count satisfying a predetermined condition, the computing unit starts other recording, copying and calculation operations, and the second count is cleared to zero and the third count is cleared to zero.
13. The processing device according to claim 12, wherein: The predetermined condition includes at least one of the following: the sum of the second count and the third count is equal to the size of the memory corresponding to each of the multiple computing units; or the sum of the second count and the third count is equal to the size of the data performed by each computing unit.
14. The processing device according to claim 12, wherein: Each computing unit increases the second count by 1 when receiving a bit related to the first tensor, and each computing unit increases the third count by 1 when receiving a bit related to the second tensor.
15. The processing device according to claim 6 or 7, wherein: The operation comprises a matrix multiplication operation, wherein the first tensor is a left matrix and the second tensor is a right matrix.
16. The processing device according to claim 6 or 7, wherein: The operation comprises a convolution operation, wherein the first tensor is an input tensor and the second tensor is a weight tensor.
17. The processing device according to claim 1, wherein: The loading comprises loading from global memory, and the copying comprises multicasting.
18. A processing method for a processing device, wherein the processing device comprises a plurality of computing units and a plurality of memories corresponding to the plurality of computing units, in, The first tensor is split into N first tensor parts, each of which is split into M segments, and the second tensor is split into M second tensor parts, each of which is split into N segments, where N and M are positive integers; The plurality of computing units at least comprises an array of N rows and M columns of computing units, The processing method comprises: The M computing units in the i-th row of the array of N rows*M columns of computing units respectively load one of the M sections in the i-th first tensor part into the memory corresponding to themselves, and copy one of the M sections in the i-th first tensor part to the memory corresponding to the remaining computing units in the i-th row of M computing units except themselves, so that the memories corresponding to the M computing units in the i-th row all have the i-th first tensor part, where i is an integer in the range [1, N], The N computing units in the j-th column of the array of N rows*M columns of computing units respectively load one segment of the N segments in the j-th second tensor part into the memory corresponding to themselves, and copy one segment of the N segments in the j-th second tensor part to the memories corresponding to the remaining computing units among the N computing units in the j-th column except themselves, so that the memories corresponding to the N computing units in the j-th column all have the j-th second tensor part, where j is an integer in the range [1, M].
19. The processing method according to claim 18, wherein: Loading, by a first computing unit (CU0) of the plurality of computing units, a first section (a0) of a first tensor portion (A0) of a first tensor into a first memory (M0) corresponding to the first computing unit (CU0); A first section (a0) of the first tensor portion (A0) is copied by the first computing unit (CU0) to a second memory (M1) corresponding to a second computing unit (CU1) of the plurality of computing units.
20. The processing method according to claim 19, wherein: The second computing unit (CU1) loads the second section (a1) of the first tensor part (A0) into the second memory (M1) and copies the second section (a1) of the first tensor part (A0) into the first memory (M0).
21. An electronic device comprising means for executing the steps of the method according to any one of claims 18 to 20.
22. A non-transitory computer readable storage medium having computer instructions stored thereon, in, When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 18-20.
23. A computer program product comprising computer instructions, in, When the computer instructions are executed by a processor, the processor is caused to perform the method according to any one of claims 18-20.
Citation Information
Patent Citations
Memory subsystem operations with unaligned and scatter gather feature to support convolution and dimension shuffle
US20190042092A1