Method for determining segmentation method, determination device, computing system, and program

The method and device determine optimal parallelization axes in hierarchical memory SIMD computers by calculating data transfer times, addressing the challenge of selecting efficient decomposition strategies and improving processing speed.

JP7715557B2Active Publication Date: 2025-07-30PREFERRED NETWORKS INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2021117748
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-07-17
Filing Date
2021-07-16
Publication Date
2025-07-30
Estimated Expiration
2041-07-16

AI Technical Summary

Technical Problem

Determining the optimal decomposition strategy for parallel computations on hierarchical memory SIMD computers is challenging due to the numerous possible strategies, making it difficult to select the most efficient approach.

Method used

A method and device that calculate data transfer times for each combination of parallelization axes in a hierarchical memory computer, considering the transfer method, problem size, and communication bandwidth, to determine the optimal combination of axes for minimizing data transfer time.

Benefits of technology

Automatically selects the optimal parallelization axes, enhancing processing speed by reducing data transfer time and improving computational efficiency in hierarchical memory systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007715557000001
    Figure 0007715557000001
  • Figure 0007715557000002
    Figure 0007715557000002
  • Figure 0007715557000003
    Figure 0007715557000003
Patent Text Reader

Abstract

To solve the problem in which it is not easy to determine an optimal policy since there are various conceivable division policies when implementing calculation that can be parallelized by recursive division on a hierarchical memory type SIMD computer.SOLUTION: One aspect of the present disclosure relates to a division system determination method that includes: a step in which one or more processors calculate data related to a data transfer time based on a transfer system determined from parallelization axes, the size of a problem to be calculated, and communication bandwidth between layers, about each combination of the parallelization axes in each layer of a hierarchical memory type computer; and a step in which the one or more processors determine the combination of specific parallelization axes based on the data related to the data transfer time calculated for each combination of the parallelization axes.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a method, a device, a computing system, and a program for determining a division scheme. [Background technology]

[0002] A parallel computer can execute computations quickly and efficiently by using multiple processors simultaneously. In parallel computing, the problem to be computed is divided into small tasks, which are then processed in parallel by each processor. A specific example of a parallel computer is a hierarchical memory SIMD (Single Instruction Multiple Data) parallel computer with a hierarchical memory architecture, which is used in supercomputers. Such computers have a hierarchical structure consisting of a total of N+1 layers, from the top layer memory (LN memory) that stores the data to be computed to the bottom layer memory (L0 memory) directly connected to the main computing units. [Prior art documents] [Non-patent literature]

[0003] [Non-Patent Document 1] Detailed explanation of GRAPE-DR, Junichiro Makino, National Astronomical Observatory of Japan Theoretical Research Division / Astronomical Simulation Project (http: / / jun.artcompsci.org / talks / tsukuba20090225.pdf) Summary of the Invention [Problem to be solved by the invention]

[0004] When implementing a computation that can be parallelized by recursive decomposition on a hierarchical memory SIMD computer, there are many possible decomposition strategies, and it is not easy to determine the optimal one. [Means for solving the problem]

[0005] To solve the above problems, one aspect of the present disclosure is that one or more processors calculate data regarding data transfer time for each combination of parallelization axes in each layer of a hierarchical memory type computer based on the transfer method determined by the parallelization axes, the size of the problem to be calculated, and the communication bandwidth between layers, and the one or more processors determine a specific combination of parallelization axes based on the data regarding the data transfer time calculated for each combination of the parallelization axes. The present disclosure relates to a method for determining a partitioning method having these steps.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Mode for Carrying Out the Invention

[0007] Hereinafter, embodiments of the present disclosure will be described. In the following examples, a determination method and a determination device for a partitioning method in a computer having a hierarchical memory architecture such as a hierarchical memory type SIMD parallel computer are disclosed. [Outline of the Present Disclosure] Outlining the present disclosure, a determination device (hereinafter referred to as a full parallelization axis determination device) calculates data regarding data transfer time based on the data transfer method determined by the parallelization axis, the size of the problem, and the communication bandwidth between layers for each combination of parallelization axes in each layer of a hierarchical memory type computer for problems that can be parallelized by recursive partitioning.

[0008] The parallelization axis means the way of partitioning in a problem that can be parallelized by recursive partitioning. For example, when the problem is the calculation of a dense matrix product A×B = C, the partitioning into partial matrix products for parallelization can be done by partitioning A in the column direction and by partitioning A in the row direction, and each of these is a parallelization axis. In a problem that can be parallelized by recursive partitioning, the same parallelization axis also exists in the problem after partitioning. When performing parallelization on a hierarchical memory type computer, it is necessary to determine the parallelization axis in each layer.

[0009] Determining the parallelization axis means determining the data transfer method between layers. If the computer has a corresponding data transfer method between the layers of interest, a partitioning method along that parallelization axis can be implemented. As data transfer methods typically provided by a hierarchical memory type computer in each layer, there are broadcast, distribute, reduce, and combine. In the above-mentioned example of a dense matrix product, the selection of two parallelization axes can be realized by a combination of these four data transfer methods.

[0010] The main arithmetic unit of the hierarchical memory type computer is directly connected to the L0 memory, where most of the operations are performed. Assuming that the total amount of operations does not change significantly with the selection of the parallelization axis, the difference in processing time for each selection is mainly determined by the difference in data transfer time between the LN memory and the L0 memory. Therefore, the combination of parallelization axes that minimizes this data transfer time can be the optimal one for increasing the processing speed of the program. Even when the amount of operations changes as a result of axis selection, a combination that minimizes the sum of the data transfer time and the calculation time can be selected, and a full parallelization axis determination device can be constructed with a simple extension. For example, consider the case where the memory of the hierarchical memory type computer is composed of three layers, namely, the L0 memory of the lowest layer directly connected to the main arithmetic unit, the L1 memory of the next layer, and the L2 memory of the highest layer. First, the full parallelization axis determination device selects from all the patterns of combinations of parallelization axes, that is, for each full parallelization axis, based on the problem size in each layer derived from the size of the problem to be calculated (for example, in the problem of obtaining the dense matrix product, the number of elements in each dimension of the matrix) and the communication bandwidth between each layer, it calculates the data transfer time between the L0 memory and the L1 memory and the data transfer time between the L1 memory and the L2 memory. Then, the largest of the calculated data transfer times is determined as the data transfer time of the corresponding full parallelization axis (step 1). Next, among the data transfer times calculated for each combination of data transfer methods, the combination of parallelization axes with the minimum data transfer time is determined as the optimal full parallelization axis (step 2). For example, in this specific example, if the data transfer time of the full parallelization axis where the parallelization axis in the L0 memory is the row direction and the parallelization axis in the L1 memory is the column direction is the minimum among all the full parallelization axes, the full parallelization axis determination device determines the full parallelization axis with the minimum data transfer time as the optimal full parallelization axis. Here, "all full parallelization axes" refers to the entire pattern of combinations of full parallelization axes such as "row direction or column direction" being the "parallelization axis", one combination such as "row direction in layer 0, row direction in layer 1, ···" being the "full parallelization axis", and combinations like "row direction in layer 0, row direction in layer 1, ···, or row direction in layer 0, column direction in layer 1 ···".What needs to be calculated is the best "fully parallelizable axis", and for this purpose, it is necessary to consider "all fully parallelizable axes". The reason why "all" appears twice is that we are discussing the combination of combinations.

[0011] According to the fully parallelizable axis determination device of the present disclosure, it is possible to automatically select the parallelizable axis in each memory layer of a hierarchical memory type computer, which was conventionally performed manually. [Computing system] First, referring to FIG. 1, a computing system according to an embodiment of the present disclosure will be described. FIG. 1 is a schematic diagram showing a computing system according to an embodiment of the present disclosure. As shown in FIG. 1, the computing system 10 includes a hierarchical memory type computer 50, an external computer 60, and a fully parallelizable axis determination device 100.

[0012] The hierarchical memory type computer 50 is composed of a hierarchical memory from L1 memory to LN memory, and a plurality of processing elements (PEs) including an arithmetic unit and L0 memory. For example, it may be a hierarchical memory type SIMD parallel computer. Further, the hierarchical memory type computer 50 may have a communication interface (IF) for communicating with the external computer 60. The communication IF may be communicatively connected to the LN memory of the top layer without being limited thereto.

[0013] For example, when the hierarchical memory type computer 50 is composed of three memory layers of L0 layer, L1 layer, and L2 layer, the L2 memory may be realized by DRAM (Dynamic Random Access Memory), the L1 memory and L0 memory may be realized by SRAM (Static Random Access Memory), and the arithmetic unit may be realized by a double-precision inner product arithmetic unit. Here, when the problem to be calculated is the calculation of a dense matrix product (A×B = C), the implementation of the problem means reading matrices A and B from DRAM to the PE and writing the calculation result C in the PE back to DRAM. Note that some operations such as reduction may be performed during the process of writing back to the top-level memory.

[0014] The external computer 60 is a computer outside the hierarchical memory type computer 50 and is communicatively connected to the hierarchical memory type computer 50. The external computer 60 may be any computer, and may be the same or similar to the hierarchical memory type computer 50.

[0015] The fully parallelized axis determination device 100 is communicatively connected to the hierarchical memory type computer 50 and determines a parallelized axis indicating a combination of memory layers in the hierarchical memory type computer 50, as will be described later.

[0016] In the illustrated embodiment, the fully parallelized axis determination device 100 is realized as an external device independent of the hierarchical memory type computer 50. However, the fully parallelized axis determination device 100 according to the present disclosure is not limited thereto and may be mounted in the hierarchical memory type computer 50. Further, the fully parallelized axis determination device 100 may be further communicatively connected to the external computer 60 and determine the parallelized axis of the hierarchical memory not only in the hierarchical memory type computer 50 but also in the external computer 60. [Fully Parallelized Axis Determination Device] Next, with reference to FIGS. 2 to 9, the fully parallelized axis determination device 100 according to an embodiment of the present disclosure will be described. FIG. 2 is a block diagram showing the functional configuration of the fully parallelized axis determination device 100 according to an embodiment of the present disclosure. As shown in FIG. 2, the fully parallelized axis determination device 100 includes a data transfer time calculation unit 110 and an optimal fully parallelized axis determination unit 120.

[0017] The data transfer time calculation unit 110 calculates the data transfer time for each combination of parallelized axes in each layer of the hierarchical memory type computer 50 based on the data transfer method determined by the parallelized axis, the size of the problem to be calculated, and the communication bandwidth between the layers. Note that the calculated data transfer time is an example of data related to the data transfer time.

[0018] For example, the data transfer method may include broadcast, distribution, reduction, and combination. Broadcast, as shown in FIG. 3, is a data transfer method that transfers the same data from a source memory to all destination memories connected to the source memory. Distribution, as shown in FIG. 4, is a data transfer method that transfers equally divided data from a source memory to all destination memories connected to the source memory. Reduction, as shown in FIG. 5, is a data transfer method that reads data of the same size from all source memories connected to a destination memory, executes any operation (e.g., sum, etc.), and transfers the result of the operation to the destination memory. Combination, as shown in FIG. 6, is a data transfer method that reads data of the same size from all source memories connected to a destination memory, combines the read data, and transfers the combined result to the destination memory.

[0019] However, the data transfer method according to the present disclosure is not limited to broadcast, distribution, reduction, and combination, and any other appropriate data transfer method may be used. Such a data transfer method can be used if there is a hardware implementation in each layer.

[0020] The data transfer time calculation unit 110 calculates the data transfer time between layers for all combinations of parallelization axes in the memory layers in the hierarchical memory type computer 50. For example, if the hierarchical memory type computer 50 is composed of three memory layers L0, L1, and L2, and the problem to be calculated has two parallelization axes, then for L1 and L0, two parallelization axes are selected respectively, so the total number of parallelization axes is 2^2 = 4. The data transfer time calculation unit 110 calculates the data transfer time for all of these 4 parallelization axes.

[0021] When the fully parallelized axis is determined, the data transfer time calculation unit 110 can determine the problem size of each layer for the problem to be calculated. For example, when the problem to be calculated is to obtain the matrix product (A×B = C), there are two parallelized axes in the column direction and the row direction, and the problem size is a set of the number of rows of A (= the number of rows of C), the number of columns of A (= the number of rows of B), and the number of columns of B (= the number of columns of C).

[0022] In the parallelization in the column direction, as shown in FIG. 7, matrix A is decomposed in the column direction and matrix B is decomposed in the row direction, and a plurality of matrices obtained by multiplying each partial matrix subA and subB are added to obtain the matrix product C. In this case, distribution is used to decompose matrices A and B, and reduction can be used to add the plurality of matrices obtained by multiplying each partial matrix. When parallelization in the column direction is performed, the number of columns of A becomes smaller in the next layer.

[0023] In the parallelization in the row direction, as shown in FIG. 8, matrix A is decomposed in the row direction, and a plurality of partial matrices obtained by multiplying each partial matrix by matrix B are combined to obtain the matrix product C. In this case, distribution is used to decompose matrix A, broadcasting is used for matrix B, and combination can be used to combine the plurality of matrices obtained by multiplying each partial matrix. When parallelization in the row direction is performed, the number of rows of A becomes smaller in the next layer.

[0024] The optimal full-parallelization axis determination unit 120 determines a specific combination of parallelization axes based on the data transfer times calculated for each combination of parallelization axes. For example, by combining the two types of parallelization axes described above, it is possible to calculate the matrix product as shown in FIG. 9. In the illustrated embodiment, the memory layer is composed of four layers, and the parallelization axis is independently determined for each layer. Specifically, the parallelization axes of layer 2, layer 1, and layer 0 are the row direction, the row direction, and the column direction, respectively. For the corresponding data transfer method, for matrix A, the data transfer methods from layer 3 to layer 2, from layer 2 to layer 1, and from layer 1 to layer 0 are all distributive. For matrix B, the data transfer methods from layer 3 to layer 2, from layer 2 to layer 1, and from layer 1 to layer 0 are broadcast, broadcast, and distributive, respectively. For matrix product C, the data transfer methods from layer 0 to layer 1, from layer 1 to layer 2, and from layer 2 to layer 3 are reduction by sum, combination, and combination, respectively.

[0025] As shown in the figure, in each layer, matrices A and B are transferred while being divided into submatrices from layer 3 to layer 0 according to the parallelization axis. The calculation result in layer 0 is aggregated through the data transfer methods from layer 0 to layer 3, and matrix product C is obtained in layer 3.

[0026] When a parallelization axis is given for a problem (such as the calculation of matrix product) for which such a divide-and-conquer method is effective, when the parallelization axis for each layer is determined, the problem size for each layer can be recursively determined, and thus the data transfer time can be estimated. When the relationship between the selection of the parallelization axis and the data transfer time is known, the optimal full-parallelization axis determination unit 120 can automatically determine the parallelization axis that minimizes the data transfer time, and this full-parallelization axis is considered to be optimal. [Parallelization Axis Determination Process] Next, with reference to FIGS. 10 and 11, a parallelization axis determination process for implementing a splitting method according to an embodiment of the present disclosure will be described. The parallelization axis determination process can be executed by the above-described all-parallelization axis determination device 100, and more specifically, by a processor of the all-parallelization axis determination device 100. FIG. 10 is a flowchart showing a parallelization axis determination process according to an embodiment of the present disclosure.

[0027] As shown in FIG. 10, in step S101, the all-parallelization axis determination device 100 calculates a data transfer time for each combination of parallelization axes in each layer of the hierarchical memory type computer 50 based on the size of the problem to be calculated and the communication bandwidth between layers. Specifically, the all-parallelization axis determination device 100 calculates the data transfer time in the data transfer method determined by the parallelization axis for each parallelization axis for each memory layer of the hierarchical memory type computer 50. For example, the all-parallelization axis determination device 100 calculates the data transfer time in the downstream direction from the top layer (LN layer) to the bottom layer (L0 layer) and the data transfer time in the upstream direction from the bottom layer (L0 layer) to the top layer (LN layer) for each combination of parallelization axes, and based on these data transfer times in the downstream direction and the upstream direction, for example, determines the data transfer time of each combination of parallelization axes as the sum of the data transfer time in the downstream direction and the data transfer time in the upstream direction.

[0028] Note that when a specific data transfer method cannot be used due to the constraints of the hierarchical memory type computer 50, combinations including parallelization axes that require the unusable data transfer method may be excluded.

[0029] In step S102, the all-parallelization axis determination device 100 determines, as the optimal all-parallelization axis, the combination of parallelization axes having the minimum data transfer time among the data transfer times calculated for each combination of parallelization axes. That is, among all possible combinations of parallelization axes calculated in step S101, the all-parallelization axis determination device 100 identifies the combination of parallelization axes that realizes the minimum data transfer time and determines the identified combination of parallelization axes as the optimal all-parallelization axis.

[0030] Figure 11 is pseudo-code showing a parallelization axis determination procedure according to an embodiment of the present disclosure. The parallelization axis determination procedure is executed by the full parallelization axis determination device 100 for the hierarchical memory type computer 50 composed of four memory layers. In this embodiment, for the problem of performing dense matrix multiplication, for each layer except L3, either the column direction or the row direction parallelization axis is independently selected, and the combination of parallelization axes that realizes the minimum data transfer time is determined as the optimal full parallelization axis. The number of selection patterns of the full parallelization axis is (the number of parallelization axis candidates)^(the number of layers - 1). In this example, it is 2^(4 - 1) = 8. It is also assumed to be sufficiently small in the actual example, and the determination of the optimal full parallelization axis can be performed by exhaustive search.

[0031] First, layers 0 to 3 are the same as the embodiment shown in FIG. 9, where layer 3 is the top layer and layer 0 corresponds to the bottom layer in the PE.

[0032] Next, L[n][m], n, m ∈ {0, 1, 2, 3} is a character string intended to represent the connection from layer n to layer m, and modifies operations with directivity between layers such as data transfer. As an exception, the parallelization axis set between each layer has no directivity, but this character string is used for convenience.

[0033] The variable transfer_time represents the minimum data transfer time among the sum of the maximum value in the downstream direction of the data transfer time between each layer from layer 3 to layer 0 and the maximum value in the upstream direction of the data transfer time between each layer from layer 0 to layer 3. Initially, the variable transfer_time is set to infinity.

[0034] The variable axes_set represents the combination of parallelization axes selected between each layer when the variable transfer_time is updated. Initially, the variable axes_set is set to (nil, nil, nil).

[0035] The constant Axes represents candidates for parallelization axes, and in this embodiment, Axes = {Column (column direction), Row (row direction)} is set.

[0036] The variable aL[i][i - 1] ∈ Axes, i ∈ {1, 2, 3} represents the parallelization axis when dividing the problem from layer i to layer i - 1.

[0037] The variables N0, N1, N2, N3 represent the problem sizes in layers 0, 1, 2, 3 respectively. The variable N[i] is represented by a tuple of one or more element counts, for example. In the case of matrix multiplication, it is represented by a tuple of two element counts. Also, a magnitude relationship is defined among the problem sizes. For example, in matrix multiplication, when N[i] = (m[i], n[i], k[i]), i ∈ {0, 1, 2, 3}, if m2 > m1 and n2 > n1 and k2 > k1, then N2 > N1.

[0038] The function Partition(N, axis, layers) returns the problem sizes when N is divided along the parallelization axis specified by axis. layers is an argument that specifies between which layers the focus of the division is, and it sets the parameters necessary for the division. In the pseudocode, it is used to determine the problem size of layer i - 1 from the problem size of layer i and the parallelization axis aL[i][i - 1], and N0 is calculated inductively from N3. Incidentally, when N[i] and N[i - 1] have a one-to-one correspondence with the parallelization axis and various parameters fixed, it is also theoretically possible to calculate the problem size of the upper layer from the problem size of the lower layer. That is, when any one of all the parallelization axes and the problem sizes N0, N1, N2, N3 is determined, the problem sizes of the remaining layers are also uniquely determined. In FIG. 11, the problem size N3 is fixed (i.e., the operation procedure is fixed) and the parallelization axis is searched. Note that when N3 ≥ N' for the originally desired problem size N', the problem can be calculated. If N3 < N', a method of dividing the originally desired problem size N' and calculating it in multiple steps can be considered.

[0039] Here, for layer 0, a lower limit value of the problem size that can be efficiently calculated by the arithmetic unit can be considered. For example, the number of elements that can be processed at once by the inner product arithmetic unit, the number of parallel executions required to hide the latency of the arithmetic unit and L0 memory, etc. affect the lower limit value. On the other hand, from the perspective of memory capacity, an upper limit value of the problem size that can be calculated can also be considered. For example, in the calculation of matrix multiplication, if the matrices A0, B0 to be multiplied by one PE and the resulting matrix C0 are stored in L0 memory, the maximum value that N0 can take can be calculated from the size of L0 memory. The upper limit value of the problem size due to memory capacity can be considered in the same way for layers other than layer 0.

[0040] The function ToSizeD(N) calculates the amount of data required for the problem represented by N in the downstream direction from layer 3 to layer 0. On the other hand, the function ToSizeU(N) calculates the amount of data required for the problem represented by N in the upstream direction from layer 0 to layer 3. In the example of matrix multiplication, since the data transferred in the downstream direction is matrices A and B, ToSizeD(N1) = m1×k1 + k1×n1, and since the data transferred in the upstream direction is matrix C, ToSizeU(N1) = m1×n1.

[0041] The function GetBW(axis, layers) calculates the communication bandwidth when axis is selected for data transfer between the layers specified by layers.

[0042] The variable tL30 represents the data transfer time from layer 3 to layer 0 and is calculated from the problem size and the communication bandwidth. On the other hand, the variable tL03 represents the data transfer time from layer 0 to layer 3. In a hierarchical memory type computer, since layer - to - layer communication can often be overlapped, the layer - to - layer communication that takes the most time becomes the bottleneck. The overlap of layer - to - layer communication means, for example, starting the communication from layer 2 to layer 1 before the communication from layer 3 to layer 2 is completed. Utilizing this, in the pseudo - code for this calculation, the maximum value max among all layer - to - layer communications is adopted. That is, based on the maximum value of the calculated data transfer time required for layer - to - layer communication, the data transfer time for each combination of parallelization axes is calculated. In a hierarchical memory type computer where layer - to - layer communication cannot be overlapped, the sum of all layer - to - layer communication times may be adopted.

[0043] As can be understood from the illustrated pseudo - code, for all combinations of the parallelization axes of the layers, first, from a given problem size N3, the problem sizes N2, N1, N0 of each layer are sequentially determined. Then, the downstream data transfer time tL30 and the upstream data transfer time tL03 are determined based on the problem size and the communication bandwidth in each layer. When the sum of the data transfer times is smaller than the current transfer_time, the sum of the data transfer times is set as the new transfer_time, and the combination of the parallelization axes is set in axes_set. This process is repeated for all combinations of the parallelization axes, and finally, the combination set in axes_set obtained is the optimal full parallelization axis.

[0044] For example, in the embodiment shown in FIG. 9, aL32, aL21, and aL10 are Row, Row, and Column respectively. Since the problem in layer 3 is the product of a 16×4 matrix and a 4×4 matrix, the problem size N3 is represented as (16, 4, 4). Next, in N2 = Partition(N3, aL32, "L32"), since the problem in layer 3 is divided by Row (row - direction division), the problem in layer 2 is the product of a 4×4 matrix and a 4×4 matrix, and N2 is represented as (4, 4, 4). Similarly, the problem sizes N1 and N0 are represented as (1, 4, 4) and (1, 1, 4) respectively. Note that only the division to N0 is by Column. Also, the communication bandwidth between each layer is determined from the chip specifications. In this way, the downstream data transfer time tL30 and the upstream data transfer time tL03 are determined based on the problem sizes and communication bandwidths in each layer. [Simulation Results] Next, referring to FIG. 12, the simulation results according to an embodiment of the present disclosure will be described. FIG. 12 is a diagram showing the simulation results according to an embodiment of the present disclosure.

[0045] In the graph of FIG. 12, d(0|1)(0|1)(0|1)m.dat indicates that the left - most (0|1) represents whether the parallelization axis of layer 2 is in the column direction or the row direction, the middle (0|1) represents whether the parallelization axis of layer 1 is in the column direction or the row direction, and the right - most (0|1) represents whether the parallelization axis of layer 0 is in the column direction or the row direction.

[0046] When the size M of the matrix is large, it is most efficient to use all Column (column - direction division) between each layer. On the other hand, when the size M of the matrix is small, it can be seen that it is efficient to use only Row (row - direction division) for data transfer between layer 3 and layer 2. Also, depending on the combination of transfer methods, there are some cases where the estimated maximum efficiency is extremely low, which is due to the combination of data transfer methods in which a sub - matrix of sufficient size is not given to the multiple PEs existing in the entire computer.

[0047] The above-described embodiments have been described by focusing on the double-precision dense matrix multiplication (DGEMM) A×B = C. However, the present disclosure is not limited thereto, and for example, it is also applicable to C = αA×B, C+ = A×B, single-precision (SGEMM), half-precision (HGEMM), and other level 3 BLAS operations. Further, the present disclosure is not limited to the above-described dense matrix calculations, and is also applicable to fully-connected layers and convolutional layers in neural networks. Further, although the procedure of rewriting the calculation result to the top layer has been described, the present disclosure is not limited thereto, and it is also applicable when the calculation result obtained in any lower layer is reused as input data for another calculation. For example, in the convolutional layer in the above-described neural network, the batch direction, the spatial direction, and the channel direction can be cited as candidates for the parallelization axis. [Hardware Configuration of All-Parallelization Axis Determination Device] In the all-parallelization axis determination device 100 according to the embodiment, each function may be a circuit configured by an analog circuit, a digital circuit, or an analog-digital hybrid circuit. Further, a control circuit for controlling each function may be provided. The implementation of each circuit may be by an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or the like.

[0048] In all of the above descriptions, at least a part of the all-parallelization axis determination device 100 may be configured by hardware, or may be configured by software, and a CPU (Central Processing Unit) or the like may perform the implementation by information processing of the software. When configured by software, a program for realizing the all-parallelization axis determination device 100 and at least a part of its functions may be stored in a non-temporary computer-readable storage medium and read and executed by a computer. The storage medium is not limited to removable ones such as magnetic disks (e.g., flexible disks) and optical disks (e.g., CD-ROMs and DVD-ROMs), and may be a fixed storage medium such as a hard disk device or an SSD (Solid State Drive) that uses a memory. That is, the information processing by software may be specifically implemented using hardware resources. Further, the processing by software may be implemented in a circuit such as an FPGA and executed by hardware. The execution of the job may be performed using an accelerator such as a GPU (Graphics Processing Unit), for example.

[0049] For example, by a computer reading a dedicated software stored in a computer-readable storage medium, the computer can be made into the device of the above embodiment. The type of the storage medium is not particularly limited. Also, by a computer installing a dedicated software downloaded via a communication network, the computer can be made into the device of the above embodiment. In this way, the information processing by software is specifically implemented using hardware resources.

[0050] FIG. 13 is a block diagram showing an example of the hardware configuration of the all-parallelization axis determination device 100 in an embodiment of the present disclosure. The all-parallelization axis determination device 100 includes a processor 101, a main storage device 102, an auxiliary storage device 103, a network interface 104, and a device interface 105, and can be realized as a computer device connected via a bus 106.

[0051] Note that although the full parallelization axis determination device 100 in FIG. 13 includes one of each component, it may include a plurality of the same components. Also, although one full parallelization axis determination device 100 is shown, software may be installed in a plurality of computer devices, and each of the plurality of full parallelization axis determination devices 100 may execute a different part of the software processing. In this case, the plurality of full parallelization axis determination devices 100 may communicate with each other via a network interface 104 or the like.

[0052] The processor 101 is an electronic circuit (processing circuit, Processing Circuit, Processing Circuitry) including a control unit and an arithmetic unit of the full parallelization axis determination device 100. The processor 101 performs arithmetic processing based on data and programs input from each device of the internal configuration of the full parallelization axis determination device 100, and outputs arithmetic results and control signals to each device and the like. Specifically, the processor 101 controls each component constituting the full parallelization axis determination device 100 by executing the OS (Operating System) of the full parallelization axis determination device 100, applications, and the like. The processor 101 is not particularly limited as long as it can perform the above processing. The full parallelization axis determination device 100 and its respective components are realized by the processor 101. Here, the processing circuit may refer to one or a plurality of electric circuits arranged on one chip, or may refer to one or a plurality of electric circuits arranged on two or more chips or devices. When using a plurality of electronic circuits, each electronic circuit may communicate with each other by wire or wirelessly.

[0053] The main memory device 102 is a storage device that stores instructions executed by the processor 101 and various data, etc., and the information stored in the main memory device 102 is directly read by the processor 101. The auxiliary storage device 103 is a storage device other than the main memory device 102. Note that these storage devices mean any electronic components capable of storing electronic information, and may be either a memory or a storage. Also, the memory includes a volatile memory and a non-volatile memory, and either may be used. The memory for storing various data within the parallelization axis determination device 100, for example, the memory may be realized by the main memory device 102 or the auxiliary storage device 103. For example, at least a part of the memory may be mounted on this main memory device 102 or the auxiliary storage device 103. As another example, when an accelerator is provided, at least a part of the aforementioned memory may be mounted in the memory provided in the accelerator.

[0054] The network interface 104 is an interface for connecting to the communication network 200 wirelessly or by wire. The network interface 104 may use one that conforms to an existing communication standard. Information exchange may be performed between the network interface 104 and an external device 109A connected via the communication network 108.

[0055] The external device 109A includes, for example, a camera, a motion capture device, an output destination device, an external sensor, an input source device, etc. Also, the external device 109A may be a device having some of the functions of the components of the entire parallelization axis determination device 100. And the parallelization axis determination device 100 may receive a part of the processing result of the entire parallelization axis determination device 100 via the communication network 108 like a cloud service.

[0056] The device interface 105 is an interface such as a USB (Universal Serial Bus) that directly connects to the external device 109B. The external device 109B may be an external storage medium or a storage device. The memory may be realized by the external device 109B.

[0057] The external device 109B may be an output device. The output device may be, for example, a display device for displaying an image, or a device for outputting sound or the like. For example, there are an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), an organic EL (ElectroLuminescence) display, a speaker, etc., but it is not limited thereto.

[0058] Note that the external device 109B may be an input device. The input device includes devices such as a keyboard, a mouse, a touch panel, and a microphone, and provides the information input by these devices to the fully parallelized axis determination device 100. The signal from the input device is output to the processor 101.

[0059] For example, the data transfer time calculation unit 110 and the optimal fully parallelized axis determination unit 120, etc. of the fully parallelized axis determination device 100 in the present embodiment may be realized by the processor 101. Further, the memory of the parallelized axis determination device 100 may be realized by the main storage device 102 or the auxiliary storage device 103. Further, the fully parallelized axis determination device 100 may be equipped with one or a plurality of memories.

[0060] In this specification, the expression "at least one of a, b, and c" or "at least one of a, b, or c" includes any combination of a, b, c, a - b, a - c, b - c, a - b - c. It also covers combinations with multiple instances of any of the elements such as a - a, a - b - b, a - a - b - b - c - c. It further covers adding other elements such as having a - b - c - d in addition to a, b, and / or c.

[0061] As described above, the embodiments of the present disclosure have been described in detail, but the present disclosure is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of the present disclosure described in the claims.

Explanation of Reference Numerals

[0062] 10 Computing system 50 Hierarchical memory type computer 60 External computer 100 Fully parallel axis determination device 110 Data transfer time calculation unit 120 Optimal fully parallel axis determination unit 101 Processor 102 Main memory device 103 Auxiliary storage device 104 Network interface 105 Device interface 108 Communication network 109 External device

Claims

1. A processing method by a computing system comprising at least one processor and a hierarchical memory type computer including three or more hierarchical memory layers, wherein: The at least one processor calculates data regarding a first data transfer time for each of a plurality of fully parallelized axes, and each of the plurality of fully parallelized axes determines a method of dividing a calculation target in each layer of the three or more memory layers. The data regarding the first data transfer time is calculated based on data regarding a plurality of second data transfer times calculated between different two of the three or more memory layers. The hierarchical memory type computer executes an operation on the calculation target based on a fully parallelized axis selected from the plurality of fully parallelized axes, and the selected fully parallelized axis is selected based on the data regarding the first data transfer time calculated for each of the plurality of fully parallelized axes. Processing method.

2. The calculation target is a matrix, and each of the plurality of fully parallelized axes determines a method of dividing the matrix in the row direction or the column direction in each layer of the three or more memory layers. The processing method according to Claim 1.

3. The three or more memory layers include at least a first layer, a second layer, and a third layer, and the data regarding the plurality of second data transfer times includes at least data regarding the data transfer time between the first layer and the second layer and data regarding the data transfer time between the second layer and the third layer. The processing method according to Claim 1 or Claim 2.

4. The data regarding the first data transfer time is the maximum value of the data regarding the plurality of second data transfer times. The processing method according to any one of Claims 1 to 3.

5. The hierarchical memory type computer is a computer capable of starting communication between another memory layer of the three or more memory layers before communication between the memory layers of the three or more memory layers is completed. The processing method according to any one of Claims 1 to 4.

6. The data regarding the first data transfer time is the sum of the data regarding the plurality of second data transfer times. The processing method according to any one of Claims 1 to 3.

7. The hierarchical memory type computer according to any one of claims 1 to 3 or claim 6 is a computer in which the start of communication between another memory layer of the three or more memory layers is impossible until the communication between the memory layers of the three or more memory layers is completed. Processing method.

8. The at least one processor determines a data transfer method between the two layers and calculates data related to the second data transfer time based on the determined data transfer method. The processing method according to any one of claims 1 to 7.

9. The object to be calculated is a layer included in a neural network. The processing method according to any one of claims 1 to 8.

10. The parallelization axis included in the fully parallelized axis includes at least any one of a batch direction, a spatial direction, or a channel direction. The processing method according to any one of claims 1 to 9.

11. A computing system comprising at least one processor and a hierarchical memory type computer including three or more hierarchical memory layers, The at least one processor calculates data related to a first data transfer time for each of a plurality of fully parallelized axes, and each of the plurality of fully parallelized axes determines a method of dividing an object to be calculated in each layer of the three or more memory layers. The data related to the first data transfer time is calculated based on data related to a plurality of second data transfer times calculated for different two layers of the three or more memory layers. The hierarchical memory type computer executes an operation on the calculation target based on a fully parallelized axis selected from the plurality of fully parallelized axes, and the selected fully parallelized axis is selected based on the data related to the first data transfer time calculated for each of the plurality of fully parallelized axes. Computing system.

12. The object to be calculated is a matrix, and each of the plurality of fully parallelized axes determines a method of dividing the row direction or the column direction of the matrix in each layer of the three or more memory layers. The computing system according to claim 11.

13. The three or more memory layers include at least a first layer, a second layer, and a third layer, and the data regarding the plurality of second data transfer times includes at least data regarding the data transfer time between the first layer and the second layer and data regarding the data transfer time between the second layer and the third layer. The computing system according to claim 11 or claim 12.

14. The data regarding the first data transfer time is the maximum value of the data regarding the plurality of second data transfer times. The computing system according to any one of claims 11 to 13.

15. The hierarchical memory type computer is a computer in which the start of communication between another memory layer of the three or more memory layers is possible before the communication between the memory layers of the three or more memory layers is completed. The computing system according to any one of claims 11 to 14.

16. The data regarding the first data transfer time is the sum of the data regarding the plurality of second data transfer times. The computing system according to any one of claims 11 to 13.

17. The hierarchical memory type computer is a computer in which the start of communication between another memory layer of the three or more memory layers is impossible until the communication between the memory layers of the three or more memory layers is completed. The computing system according to any one of claims 11 to 13 or claim 16.

18. The at least one processor determines the data transfer method between the two layers, and calculates the data regarding the second data transfer time based on the determined data transfer method. The computing system according to any one of claims 11 to 17.

19. The object to be computed is a layer included in a neural network. The computing system according to any one of claims 11 to 18.

20. The parallelization axes included in the fully parallelized axis include at least any one of the batch direction, the spatial direction, or the channel direction. The computing system according to any one of claims 11 to 19.

Citation Information

Patent Citations

  • Area partition pattern decision method

    JP2003248667A

  • Optimum code generation method for multiprocessor, and compiling device

    JP2009104422A

  • Apparatus and method for partitioning programs between a general purpose core and one or more accelerators

    US20070174828A1