A method and device for adaptive distribution of DW type operator data in a multi-core environment
By adaptively searching and adapting data distribution and optimizing data division and calculation in a multi-core environment, the problem of unreasonable arrangement of operator tasks in the AI chip core group in the prior art is solved, and more efficient computing resource utilization and operator performance optimization are achieved.
Patent Information
- Application Number
- CN202411629645.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-15
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2044-11-15
AI Technical Summary
In the current stage of operator development, it is impossible to properly arrange operator tasks in the core group based on the AI chip structure, resulting in insufficient utilization of space and computing resources, which leads to low operator performance or insignificant optimization.
A method of data distribution of dw type operators adaptively in a multi-core environment is proposed. By obtaining the parameters of hardware devices and calculation tasks, adaptively searching for adaptive data distribution, and dividing the input data into multiple blocks for calculation, selecting the specification dimension and the connection write back dimension to optimize the data distribution.
Through the adaptive data distribution method, the additional data transmission overhead caused by unreasonable data distribution is reduced, the operator performance is optimized, and the utilization rate of computing resources is improved.
Smart Images

Figure CN119166948B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the intersecting field of intelligent computing, high-performance computing, operator development, deep learning, and machine learning, and specifically relates to a method and device for adaptively distributing DW-type operator data in a multi-core environment. Background Art
[0002] Intelligent computing is a computer technology that aims to simulate or achieve capabilities similar to human intelligence. It covers multiple subfields such as machine learning, deep learning, natural language processing, computer vision, etc. The goal of intelligent computing is to imitate certain aspects of human intelligence through computer systems to achieve automatic parallel and intelligent task execution and decision-making.
[0003] Operator library is a very important part of intelligent computing. Operator can be understood as a specific computing operation, which is used to represent and process data in intelligent computing. Operator library is a set of software libraries or programming frameworks that provide various commonly used operators. These operators can be mathematical operations, data conversion, feature extraction, model optimization, and so on. Operator library provides reusable operator implementations, enabling developers to build intelligent computing systems more conveniently. By using operator library, developers can call ready-made operators in their applications without having to implement them from scratch. In this way, they can focus more on high-level model design and system development, speed up development, and reduce code complexity.
[0004] The relationship between intelligent computing and operator libraries can be seen as a function and tool relationship. Intelligent computing provides an overall idea and paradigm, while the operator library, as a part of it, provides specific tools and support for the implementation of intelligent computing. By using the operator library, developers can use existing operators to build intelligent computing systems to improve efficiency and quality.
[0005] From the perspective of intelligent computing, the challenges faced by operator library development mainly include:
[0006] 1) Complexity and diversity of operators;
[0007] 2) Operator performance and efficiency optimization;
[0008] 3) The hardware and software ecosystems are incompatible.
[0009] Among them, the inability to reasonably arrange operator tasks in the core group according to the AI chip structure leads to insufficient utilization of space and computing resources, which further leads to low operator performance or unclear optimization, which is a common problem in operator development at this stage. Due to the complexity of operators (mainly reflected in different input and output shapes of the same function), it is often necessary to manually design different solutions for different shapes of the same operator. Summary of the invention
[0010] In view of the shortcomings of the prior art, the present invention proposes an adaptive dw type operator data distribution method in a multi-core environment, and the specific technical solution is as follows:
[0011] The purpose of the present invention is achieved through the following technical solutions:
[0012] An adaptive DW type operator data distribution method in a multi-core environment, comprising:
[0013] Step 1: Obtain the hardware device parameters and computing task parameters involved in the calculation;
[0014] Step 2: searching for data distribution adapted to the hardware device parameters according to the shapes of the input feature map x of the forward propagation and the input feature map dy of the backward propagation in the computing task parameters;
[0015] Step 3: Divide the input data into multiple blocks for calculation according to the data distribution of the hardware device parameters in step 2 and the single data acquisition size;
[0016] Step 4: Select the reduction dimension and the connection write-back dimension to write back according to the data distribution obtained in step 2.
[0017] Furthermore, in the step 1, the hardware device parameters involved in the calculation include the number of computing resources included in the computing chip core group, the computing resource number, and the row and column distribution parameters of the computing resources; the computing task parameters include the dimension and shape of the input feature graphs x and dy, the size and step size of the convolution kernel, and the common cumulative dimension of x and dy, and the non-cumulative dimension of x and dy respectively;
[0018] The step 2 specifically includes first comparing the number of computing resources COL on the column of the computing resource core group with the number of computing resources ROW on the row of the computing resource core group, and then continuing to perform the following operations in both the cases of COL>ROW and COL<=ROW: first determine whether to split the channel size C dimension of the input feature map x, then determine whether to split the channel size M dimension of the input feature map dy, and finally determine whether to split the accumulation dimensions N*H*W and N*E*F*R*S; wherein N represents the batch size BatchSize, H represents the height of the input feature map x, W represents the width of the input feature map x, E represents the height of the input feature map dy, F represents the width of the input feature map dy, R represents the height of the input convolution kernel, and S represents the width of the input convolution kernel.
[0019] Furthermore, when COL>ROW, determining whether to split the C dimension specifically includes:
[0020] Judge the size of C and the row MXU_LEFT_ROW of the left matrix of the multiplication calculation unit. If C < MXU_LEFT_ROW, do not split the dimension of C; otherwise, split the dimension of C. The specific splitting strategy is as follows:
[0021] Judge the relationship between C / MXU_LEFT_ROW and ROW:
[0022] If C / MXU_LEFT_ROW <= ROW, split the dimension of C, perform column broadcasting on x, and each computing unit on each column holds x with a size of ceil(C / MXU_LEFT_ROW * N * H * W); where ceil() is the ceiling function;
[0023] If C / MXU_LEFT_ROW > ROW, further judge the relationship between C and M: If C > M, perform column broadcasting on x, and each computing unit on each column holds x with a size of ceil(C / MXU_LEFT_ROW * N * H * W); if C <= M, do not split the dimension of C; thus completing the judgment of the splitting situation of the C dimension.
[0024] Furthermore, when COL > ROW, the specific steps for judging whether to split the M dimension include:
[0025] If C * M <= max(N * H * W, N * E * F * R * S), do not split the dimension of M;
[0026] If C * M > max(N * H * W, N * E * F * R * S), then further judge the relationship between M / MXU_RT_COL and COL; where MXU_RT_COL represents the column of the right matrix of the multiplication calculation unit;
[0027] If M / MXU_RT_COL > COL, do not split the dimension of M;
[0028] If M / MXU_RT_COL <= COL, split the dimension of M, perform row broadcasting on dy, so that each computing unit on each row holds dy with a size of ceil(M / MXU_RT_COL) * N * E * F * R * S; thus completing the judgment of the splitting situation of the M dimension.
[0029] Furthermore, when COL > ROW, the sub-steps for judging whether to split the accumulation dimensions N * H * W and N * E * F * R * S specifically include:
[0030] If both the C and M dimensions are split, do not split the accumulation dimensions N * H * W and N * E * F * R * S;
[0031] If neither the C nor the M dimension is split, then split the accumulation dimensions N*H*W and N*E*F*R*S simultaneously; after splitting, each column computing unit holds an x of size N*H*W / COL and a dy of size N*E*F*R*S / COL; further judge the sizes of C and M. If C<M, broadcast x column-wise and store dy column-distributedly; otherwise, broadcast dy column-wise and store x column-distributedly.
[0032] If exactly one of the C and M dimensions is not split, then split the accumulation dimensions N*H*W and N*E*F*R*S simultaneously; after splitting, each column computing unit holds an x of size N*H*W / ROW and a dy of size N*E*F*R*S / ROW.
[0033] Furthermore, when COL <= ROW, judging whether to split the C dimension specifically includes:
[0034] Judge the size relationship between C and MXU_LEFT_ROW. If C < MXU_LEFT_ROW, then do not split the C dimension; otherwise, split the C dimension. The specific splitting strategy is as follows:
[0035] Judge the relationship between C / MXU_LEFT_ROW and COL:
[0036] If C / MXU_LEFT_ROW <= COL, then split the C dimension, broadcast x row-wise, and each computing unit on each row holds an x of size ceil(C / MXU_LEFT_ROW)*N*H*W.
[0037] If C / MXU_LEFT_ROW > COL, then further judge the relationship between C and M. If C > M, then broadcast x row-wise, and each computing unit on each row holds an x of size ceil(C / MXU_LEFT_ROW)*N*H*W; if C <= M, do not split the C dimension; thus completing the judgment of the C dimension splitting situation.
[0038] Furthermore, when COL <= ROW, judging whether to split the M dimension specifically includes:
[0039] If C*M <= max(N*H*W, N*E*F*R*S), then do not split the M dimension;
[0040] If C*M > max(N*H*W, N*E*F*R*S), then further judge the relationship between M / MXU_RT_COL and ROW; if M / MXU_RT_COL > ROW, then do not split the M dimension; if M / MXU_RT_COL <= ROW, then split the M dimension, broadcast dy column-wise, so that each computing unit on each column holds a dy of size ceil(M / MXU_RT_COL)*N*E*F*R*S; thus completing the judgment of the M dimension splitting situation.
[0041] Further, when COL <= ROW, the sub-steps of determining whether to split the accumulation dimensions N*H*W and N*E*F*R*S specifically include:
[0042] If both the C and M dimensions are split, then do not split the accumulation dimensions N*H*W and N*E*F*R*S;
[0043] If neither the C nor the M dimension is split, then split both the accumulation dimensions N*H*W and N*E*F*R*S simultaneously; after splitting, each column computing unit holds an x of size N*H*W / ROW and a dy of size N*E*F*R*S / ROW; further determine the sizes of C and M, if C < M, broadcast x row-wise and store dy in a row-distributed manner, otherwise, broadcast dy row-wise and store x in a row-distributed manner;
[0044] If exactly one of the C and M dimensions is not split, then split both the accumulation dimensions N*H*W and N*E*F*R*S simultaneously; after splitting, each column computing unit holds an x of size N*H*W / COL and a dy of size N*E*F*R*S / COL.
[0045] Further, the specific steps of step four specifically include the following sub-steps:
[0046] S4.1: Determine the broadcast form. If it is row broadcast:
[0047] (1) Further determine the object to be broadcast; if the object to be broadcast is x, the shape of the result dw obtained by each computing resource is <the non-accumulation dimensions of x, the non-accumulation dimensions of dy / COL>; perform a column reduction operation on dw, and then reduce the result to the same row. The shape of the result dw obtained by each computing resource on the row is <the non-accumulation dimensions of x, the non-accumulation dimensions of dy / COL>; if the object to be broadcast is dy, the shape of the result dw obtained by each computing resource is <the non-accumulation dimensions of x / COL, the non-accumulation dimensions of dy>. Perform a column reduction operation on dw, and then reduce the result to the same row. The shape of the result dw obtained by each computing resource on the row is <the non-accumulation dimensions of x / COL, the non-accumulation dimensions of dy>;
[0048] (2) For each result dw on the reduced row, perform a concatenation write-back operation to obtain the final shape of dw as <the non-accumulation dimensions of x, the non-accumulation dimensions of dy>;
[0049] Determine the broadcast form. If it is column broadcast:
[0050] (a) Further determine the broadcast object; if the broadcast object is x, the shape of the result dw obtained by each computing resource is <the non-accumulative dimensions of x, the non-accumulative dimensions of dy / ROW>. Perform a row reduction operation on dw, and then reduce the result to the same column. The shape of the result dw obtained by each computing resource on the column is <the non-accumulative dimensions of x, the non-accumulative dimensions of dy / ROW>; if the broadcast object is dy, the shape of the result dw obtained by each computing resource is <the non-accumulative dimensions of x / ROW, the non-accumulative dimensions of dy>. Perform a row reduction operation on dw; then reduce the result to the same column. The shape of the result dw obtained by each computing resource on the column is <the non-accumulative dimensions of x / ROW, the non-accumulative dimensions of dy>;
[0051] (b) For each result dw on the reduced column, perform a concatenation write-back operation to obtain the final shape of dw as <the non-accumulative dimensions of x, the non-accumulative dimensions of dy>.
[0052] An adaptive data distribution device for dw type operators in a many-core environment, including a parameter acquisition module, a data distribution determination module, a data partitioning and calculation module, and a data write-back module;
[0053] The parameter acquisition module is used to acquire the hardware device parameters and calculation task parameters involved in the calculation;
[0054] The data distribution determination module searches for a data distribution that adapts to the hardware device parameters according to the shapes of the input feature map x in the forward propagation and the input feature map dy in the backward propagation in the calculation task parameters;
[0055] The data partitioning and calculation module divides the input data into multiple blocks for calculation according to the data distribution of the hardware device parameters and the single fetch size;
[0056] The data write-back module writes back according to the data distribution obtained by the data distribution determination module by selecting the reduction dimension and the concatenation write-back dimension.
[0057] The beneficial effects of the present invention are as follows:
[0058] The adaptive data distribution method and device for dw type operators in the many-core environment of the present invention can adaptively search for a data distribution that adapts to the calculation for the computing chip and the data parameters involved in the calculation, thereby reducing the overhead of additional data transmission caused by unreasonable data distribution, and thus optimizing the operator performance. Description of the Drawings
[0059] Figure 1 It is a flowchart of the adaptive data distribution method for dw type operators in the many-core environment of the embodiment of the present invention.
[0060] Figure 2This is a flow chart of searching data distribution when COL>ROW according to an embodiment of the present invention.
[0061] Figure 3 This is a flow chart of searching data distribution when COL<=ROW according to an embodiment of the present invention.
[0062] Figure 4 The present invention is a flowchart of selecting a reduction dimension and connecting a write-back dimension to write back according to data distribution according to an embodiment of the present invention.
[0063] Figure 5 is the original data form of x and the searched data distribution in the embodiment of the present invention; wherein Figure 5 (a) shows the original data form of x, and (b) shows the adaptive data distribution of x after adopting the method of this embodiment.
[0064] Figure 6 is the original data form of dy and the searched data distribution in the embodiment of the present invention; wherein Figure 6 Figure (a) shows the original data form of dy, and Figure (b) shows the data distribution of dy adaptively after adopting the method of this embodiment. DETAILED DESCRIPTION
[0065] The present invention will be described in detail below based on the accompanying drawings and preferred embodiments, and the purpose and effects of the present invention will become more clear. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0066] The experimental environment of this embodiment is a single core group on a single node and a single card. The experimental object is the convolution calculation back propagation calculation weight (kernel) operator.
[0067] like Figure 1 As shown, the adaptive dw type operator data distribution method in the multi-core environment of this embodiment specifically includes the following steps:
[0068] Step 1: Obtain the hardware device parameters and computing task parameters involved in the calculation;
[0069] Among them, the hardware device parameters involved in the calculation include the number of computing resources CORE_NUM contained in the computing chip core group, the computing resource number Tid, and the row and column distribution parameters of the computing resources. The row and column distribution parameters of the computing resources specifically include the number of computing resources ROW on the rows of the computing resource core group and the number of computing resources COL on the columns of the computing resource core group. In this embodiment, their values are shown in Table 1.
[0070] Table 1 Parameters of hardware devices involved in the calculation
[0071]
[0072] The calculation task parameters involved in the calculation include the dimensions and shapes of the input feature map x and dy, the size of the convolutional kernel, the stride, as well as the cumulative dimension common to x and dy and the non-cumulative dimensions of x and dy respectively. The specific values of the calculation task parameters are shown in Table 2.
[0073] Table 2 Calculation Task Parameter Table
[0074]
[0075] Step 2: Search for the data distribution that adapts to the hardware device parameters according to the shapes of x and dy in the calculation task parameters. The operations used in this step are shown in Table 3.
[0076] Table 3 Operation Name Table
[0077]
[0078] As Figure 2 、 Figure 3 and Figure 4 shown, Step 2 specifically includes first comparing the number of computing resources COL on the computing resource kernel group column with the number of computing resources ROW on the computing resource kernel group row, and then continuing to execute the following operations in both cases of COL > ROW and COL <= ROW: First, judge whether to split the channel size C dimension of the input feature map x, then judge whether to split the channel size M dimension of the input feature map dy, and finally judge whether to split the cumulative dimension N*H*W or N*E*F*R*S.
[0079] In this embodiment, since COL > ROW (8 > 4), it is further judged whether to split the C dimension; since C < MXU_LEFT_ROW (64 < 128), the C dimension is not split; next, it is judged whether to split the M dimension. Since C * M < max(N*H*W, N*E*F*R*S) (64 * 256 < max(64 * 56 * 56, 64 * 56 * 56 * 1 * 1)), the M dimension is not split; next, it is judged whether to split the cumulative dimensions N*H*W and N*E*F*R*S. Since neither the C dimension nor the M dimension is split, both the cumulative dimensions N*H*W and N*E*F*R*S are split; after splitting, each column of computing units holds x with a size of N*H*W / COL and dy with a size of N*E*F*R*S / COL; judge the relationship between the non-cumulative dimension C and the non-cumulative dimension M. Since C < M (64 < 256), x is stored by column broadcast and dy is stored by column distribution. Thus, the distribution form of the input data x and dy on the computing resources is obtained.
[0080] It is known that the object of column broadcast is x. Then, as Figure 5 shown, Figure 5Figure (a) shows the original data form of x. Figure 5 Figure (b) in the figure shows the adaptive data distribution of x after adopting the method of this embodiment. The computing resource storage of each column accounts for 1 / 8 of the total amount of x, that is, NBN (NBN=N*E*F / 8, which means that the data of size N*E*F is divided into 8 parts, each part consists of an equal number of elements) * C, and the shape of x is [8*56*56*64], and the stride is 1605632 (8*56*56*64). The computing resources on the same column store the same amount of data. The reading of each column of data is completed by initiating a column broadcast after any computing resource on the column reads it. The stride size of 1605632 refers to the byte distance from the starting address of a data block to the starting address of the next data block at the same position in the memory. This allows each tid to read data in parallel at different memory locations to optimize the efficiency of parallel processing. For example, tid=0 starts reading data from memory address 0, while tid=17 starts reading from memory address 1605632, which means that tid=17 accesses the starting position of the next data block. Corresponding to x, each column of computing resources is distributed and stored in dy. Figure 6 As shown, Figure 6 Figure (a) shows the original data form of dy. Figure 6 Figure (b) shows the data distribution of dy after the method of this embodiment is adopted. The dy stored in each column covers the continuous part of the original data with a size of NBN*M and a shape of [25088*256]. Each computing resource stores the continuous part of the original data of dy with a size of NBN*bM (bM=M / 4, which means that the data of size M is divided into 4 parts, each part consists of an equal number of elements) and a shape of [25088*64].
[0081] Step 3: According to the data distribution of the hardware device parameters in step 2 and the single data acquisition size, divide the input data into multiple blocks for calculation. Figure 6 As shown in FIG. 2( b ), in this embodiment, the single data fetch size is set to 8192. The purpose of data block division is to read the entire data of the computing resource in a circular manner and perform calculation according to the storage space size of the computing resource.
[0082] Step 4: Select the reduction dimension and the connection write-back dimension for the data distribution obtained in step 2.
[0083] Select the Reduce dimension and Concat dimension according to the data layout; first, judge the broadcast form: column broadcast; secondly, judge the object to be broadcast: x; it can be confirmed that the shape of the result dw obtained by each computing resource is <the non-accumulative C of x, the non-accumulative dimension M / ROW of dy>; then perform a Reduce operation on dw in the row direction, and the result dw is reduced to the same column computing resource, and the shape of the result dw obtained by each computing resource on the column is <the non-accumulative C of x, the non-accumulative dimension M / ROW of dy>;
[0084] For each result dw on the reduced column, perform a concatenation write-back operation. The size of each data block is <the non-accumulative C of x, the non-accumulative dimension M / ROW of dy>, and the element cross-block length is <C, M / ROW>; after completing the concatenation write-back operation, the final shape of dw is <the non-accumulative C of x, the non-accumulative dimension M of dy>.
[0085] Those of ordinary skill in the art can understand that the above are only preferred examples of the invention and are not used to limit the invention. Although the invention has been described in detail with reference to the foregoing examples, those skilled in the art can still modify the technical solutions described in the foregoing examples, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, etc. made within the spirit and principle of the invention shall be included within the protection scope of the invention.
Claims
1. An adaptive dw type operator data distribution method in a many-core environment, characterized in that: Including: Step 1: Obtain the hardware device parameters and calculation task parameters involved in the calculation; Step 2: Search for the data distribution adapted to the hardware device parameters according to the shapes of the input feature map x in the forward propagation and the input feature map dy in the backward propagation in the calculation task parameters; Step 3: Divide the input data into multiple blocks for calculation according to the data distribution of the hardware device parameters in Step 2 and the size of each data fetch; Step 4: Select the reduction dimension and the connection write-back dimension for write-back according to the data distribution obtained in Step 2; In Step 1, the hardware device parameters involved in the calculation include the number of computing resources in the computing chip core group, the computing resource numbers, and the row and column distribution parameters of the computing resources; the calculation task parameters include the dimensions and shapes of the input feature maps x and dy, the size of the convolution kernel, the stride, as well as the common accumulation dimension of x and dy and the non-accumulation dimensions of x and dy respectively; Step 2 specifically includes first comparing the number of computing resources COL in the column of the computing resource core group with the number of computing resources ROW in the row of the computing resource core group, and then continuing to perform the following operations in both cases of COL > ROW and COL <= ROW: first judge whether to split the channel size C dimension of the input feature map x, then judge whether to split the channel size M dimension of the input feature map dy, and finally judge whether to split the accumulation dimensions N*H*W and N*E*F*R*S; where, N represents the batch size, H represents the height of the input feature map x, W represents the width of the input feature map x, E represents the height of the input feature map dy, F represents the width of the input feature map dy, R represents the height of the input convolution kernel, and S represents the width of the input convolution kernel; When COL > ROW, judging whether to split the C dimension specifically includes: Judging the relationship between C and the number of rows MXU_LEFT_ROW of the left matrix of the multiplication computing unit. If C < MXU_LEFT_ROW, then do not split the C dimension; otherwise, split the C dimension, and the specific splitting strategy is: Judging the relationship between C / MXU_LEFT_ROW and ROW: If C / MXU_LEFT_ROW <= ROW, then split the C dimension and perform column broadcast on x, and each computing unit in the column holds x with a size of ceil(C / MXU_LEFT_ROW*N*H*W); where, ceil() is the ceiling function; If C / MXU_LEFT_ROW > ROW, further judge the relationship between C and M: if C > M, then perform column broadcast on x, and each computing unit in the column holds x with a size of ceil(C / MXU_LEFT_ROW*N*H*W); if C <= M, do not split the C dimension; thus completing the judgment of the C dimension splitting situation; When COL > ROW, judging whether to split the M dimension specifically includes: If C*M <= max(N*H*W, N*E*F*R*S), then do not split the M dimension; If C * M > max(N * H * W, N * E * F * R * S), then further determine the relationship between M / MXU_RT_COL and COL; where MXU_RT_COL represents the right matrix column of the multiplication calculation unit. If M / MXU_RT_COL > COL, then do not split the M dimension. If M / MXU_RT_COL <= COL, then split the M dimension, and perform row broadcast on dy so that each calculation unit in each row holds dy with a size of ceil(M / MXU_RT_COL) * N * E * F * R * S; thus completing the determination of the M dimension splitting situation. When COL > ROW, the sub - steps for determining whether to split the accumulation dimensions N * H * W and N * E * F * R * S specifically include: If both the C and M dimensions are split, then do not split the accumulation dimensions N * H * W and N * E * F * R * S. If neither the C nor the M dimension is split, then split both the accumulation dimensions N * H * W and N * E * F * R * S simultaneously; after splitting, each column calculation unit holds x with a size of N * H * W / COL and dy with a size of N * E * F * R * S / COL; further determine the sizes of C and M. If C < M, perform column broadcast to store x and column - distributed storage of dy. Otherwise, perform column broadcast to store dy and column - distributed storage of x. If only one of the C and M dimensions is not split, then split both the accumulation dimensions N * H * W and N * E * F * R * S simultaneously; after splitting, each column calculation unit holds x with a size of N * H * W / ROW and dy with a size of N * E * F * R * S / ROW. When COL <= ROW, the determination of whether to split the C dimension specifically includes: Determine the size relationship between C and MXU_LEFT_ROW. If C < MXU_LEFT_ROW, then do not split the C dimension; otherwise, split the C dimension. The specific splitting strategy is as follows: Determine the relationship between C / MXU_LEFT_ROW and COL: If C / MXU_LEFT_ROW <= COL, then split the C dimension and perform row broadcast on x so that each calculation unit in each row holds x with a size of ceil(C / MXU_LEFT_ROW) * N * H * W. If C / MXU_LEFT_ROW > COL, then further determine the relationship between C and M. If C > M, perform row broadcast on x so that each calculation unit in each row holds x with a size of ceil(C / MXU_LEFT_ROW) * N * H * W. If C <= M, do not split the C dimension; thus completing the determination of the C dimension splitting situation. When COL <= ROW, the determination of whether to split the M dimension specifically includes: If C * M <= max(N * H * W, N * E * F * R * S), then do not split the M dimension. If C * M > max(N * H * W, N * E * F * R * S), then further judge the relationship between M / MXU_RT_COL and ROW; if M / MXU_RT_COL > ROW, do not split the M dimension; if M / MXU_RT_COL <= ROW, split the M dimension, and column-broadcast dy so that each computing unit in each column holds dy with a size of ceil(M / MXU_RT_COL) * N * E * F * R * S; thus, complete the judgment of the M-dimension splitting situation; When COL <= ROW, the sub-steps for judging whether to split the accumulation dimensions N * H * W and N * E * F * R * S specifically include: If both the C and M dimensions are split, do not split the accumulation dimensions N * H * W and N * E * F * R * S; If neither the C nor the M dimension is split, split both the accumulation dimensions N * H * W and N * E * F * R * S simultaneously; after splitting, each computing unit in each column holds x with a size of N * H * W / ROW and dy with a size of N * E * F * R * S / ROW; further judge the sizes of C and M, if C < M, row-broadcast and store x, and store dy in a row-distributed manner, otherwise, row-broadcast and store dy, and store x in a row-distributed manner; If exactly one of the C and M dimensions is not split, split both the accumulation dimensions N * H * W and N * E * F * R * S simultaneously; after splitting, each computing unit in each column holds x with a size of N * H * W / COL and dy with a size of N * E * F * R * S / COL.
2. The adaptive dw type operator data distribution method in a multi-core environment according to claim 1, characterized in that: The specific steps of Step 4 specifically include the following sub-steps: S4.1: Judge the broadcast form. If it is row-broadcast: (1) Further judge the object to be broadcast; if the object to be broadcast is x, the shape of the result dw obtained by each computing resource is <the non-accumulation dimensions of x, the non-accumulation dimensions of dy / COL>; perform a column reduction operation on dw, and then reduce the result to the same row, and the shape of the result dw obtained by each computing resource on the row is <the non-accumulation dimensions of x, the non-accumulation dimensions of dy / COL>; if the object to be broadcast is dy, the shape of the result dw obtained by each computing resource is <the non-accumulation dimensions of x / COL, the non-accumulation dimensions of dy>, perform a column reduction operation on dw, and then reduce the result to the same row, and the shape of the result dw obtained by each computing resource on the row is <the non-accumulation dimensions of x / COL, the non-accumulation dimensions of dy>; (2) Perform a concatenation write-back operation on each result dw on the reduced row to obtain the final shape of dw as <the non-accumulation dimensions of x, the non-accumulation dimensions of dy>; Judge the broadcast form. If it is column-broadcast: (a)Further determine the broadcast object; if the broadcast object is x, the shape of the result dw obtained by each computing resource is <non-accumulative dimensions of x, non-accumulative dimensions of dy / ROW>. Perform a row reduction operation on dw, and then reduce the result to the same column. The shape of the result dw obtained by each computing resource on the column is <non-accumulative dimensions of x, non-accumulative dimensions of dy / ROW>; if the broadcast object is dy, the shape of the result dw obtained by each computing resource is <non-accumulative dimensions of x / ROW, non-accumulative dimensions of dy>. Perform a row reduction operation on dw, and then reduce the result to the same column. The shape of the result dw obtained by each computing resource on the column is <non-accumulative dimensions of x / ROW, non-accumulative dimensions of dy>. (b)For each result dw on the reduced column, perform a concatenation write-back operation to obtain the final shape of dw as <non-accumulative dimensions of x, non-accumulative dimensions of dy>.
3. An adaptive dw type operator data distribution device in a multi-core environment, characterized in that: A method for implementing an adaptive data distribution of the dw type operator in a many-core environment according to claim 1, including a parameter acquisition module, a data distribution determination module, a data partitioning and calculation module, and a data write-back module; The parameter acquisition module is used to acquire the hardware device parameters and computing task parameters involved in the calculation; The data distribution determination module searches for a data distribution that adapts to the hardware device parameters according to the shapes of the input feature map x in the forward propagation and the input feature map dy in the backward propagation in the computing task parameters; The data partitioning and calculation module divides the input data into multiple blocks for calculation according to the data distribution of the hardware device parameters and the single fetch size; The data write-back module selects the reduction dimension and the concatenation write-back dimension for write-back according to the data distribution obtained by the data distribution determination module.
Citation Information
Patent Citations
Floating point matrix multiplier many-core parallel optimization method for deep learning
CN112732630A
Machine learning computation optimization method and platform
WO2023179415A1