Computation device and data transfer method
By associating external memories with processors to optimize data transfer paths and grouping processors, the computing device addresses inefficiencies in data transfer preparation times, enhancing computational efficiency through simultaneous data transfer.
Patent Information
- Application Number
- PCT/JP2025/023111
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-06-26
- Publication Date
- 2026-02-12
AI Technical Summary
Existing computing devices with multiple processors and external memories face inefficiencies in data transfer preparation times due to the need for rearranging data based on processor arrangement, leading to prolonged data transfer preparation times.
A computing device configuration where external memories are associated with processors to predetermine data transfer paths, allowing for simultaneous data transfer in multiple directions and grouping processors to facilitate efficient data rearrangement and aggregation, thereby reducing preparation time.
This configuration significantly reduces data transfer preparation time and enhances the capacity for simultaneous data transfer, improving overall computational efficiency.
Smart Images

Figure JP2025023111_12022026_PF_FP_ABST
Abstract
Description
Arithmetic device and data movement method CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is based on Japanese Application No. 2024-130515, filed on August 7, 2024, the contents of which are incorporated herein by reference.
[0002] The present disclosure relates to a computing device and a data movement method.
[0003] 2. Description of the Related Art Arithmetic units have been developed that are configured with a plurality of processors arranged in a two-dimensional configuration (for example, an array configuration, also called a mesh configuration).
[0004] As such a computing device, for example, Patent Document 1 describes a computing device including a first plurality of processing cores arranged in an array and a second plurality of processing cores arranged in an array.
[0005] U.S. Pat. No. 1,086,574
[0006] 25, a conventional arithmetic device 100 includes a plurality of processors (Processing Elements, hereinafter referred to as "PEs") 102 and an external memory 104. Note that arrows in FIG. 25 indicate examples of data input / output between the PEs 102 and between the PEs 102 and the external memory 104.
[0007] When data is transferred from one external memory 104 to multiple PEs 102, the data is rearranged in various ways, such as arranging the data horizontally or vertically depending on the arrangement of the PEs 102, or making the data the same for all PEs 102, before being transferred to the PEs 102. For this reason, it takes time to prepare for transferring data from the external memory 104 to the PEs 102.
[0008] An object of the present disclosure is to provide a computing device and a data transfer method that can shorten the preparation time for data transfer between an external memory and a processor.
[0009] A computing device according to one aspect of the present disclosure is a computing device comprising a plurality of external memories and a plurality of processors, each of which is associated with an external memory that inputs and outputs data, and each of which stores data in association with the processor that inputs and outputs data.
[0010] According to this configuration, the external memory and processor to which data is transferred are predetermined, so that the preparation time for data transfer between the external memory and the processor can be shortened.
[0011] In the above-described computing device, data transferred from an external memory provided outside the computing device to the external memory may be sorted in accordance with the location of the processor and then stored in the external memory. With this configuration, it is possible to shorten the preparation time for data transfer between the external memory and the processor.
[0012] In the computing device, the external memories may be arranged in a plurality of different directions relative to the processor. With this configuration, more data can be transferred at one time.
[0013] In the computing device, the external memory may be arranged in a one-to-one correspondence with the processor. With this configuration, more data can be transferred at one time.
[0014] In the computing device, the processors may be divided into a plurality of groups in a row or a column, and the external memory may be arranged for each group. With this configuration, more data can be transferred at one time.
[0015] In the computing device, the processors may be divided into a plurality of groups of at least one of a plurality of rows and a plurality of columns, and the external memory may be arranged for each group. With this configuration, more data can be transferred at one time.
[0016] In the above-described arithmetic device, data output from the plurality of processors arranged in the same row to the external memory may be input from the external memory to the plurality of processors arranged in the same column, and data output from the plurality of processors arranged in the same column to the external memory may be input from the external memory to the plurality of processors arranged in the same row. With this configuration, it is possible to easily interchange the rows and columns of data held by the processors.
[0017] In the above-described arithmetic device, data output from the plurality of processors to the external memory may be input from the external memory to the processor so that one of the processors holds the plurality of data. With this configuration, the plurality of data can be consolidated and held in the processor.
[0018] In the above-described arithmetic device, data output from the plurality of processors to the external memory may be copied and input from the external memory to the processors so that the plurality of processors hold the same data. According to this configuration, the plurality of data may be distributed to the processors and held therein.
[0019] In the above-mentioned arithmetic device, data formed into a multidimensional array by a plurality of groups formed of two-dimensionally arranged data may be stored in an external memory provided outside the arithmetic device, and after inputting the data in the nth row and mth column of the plurality of groups from the external memory to the processor in the nth row and mth column, the processor may output the data to the corresponding external memory, and the external memory may input the data to the plurality of processors so that the arrangement of the data formed into the multidimensional array by the plurality of groups becomes a two-dimensional array. With this configuration, the multidimensionally arranged data can be arranged in the two-dimensionally arranged processors.
[0020] In the above-mentioned arithmetic device, one or more of the external memories may be associated with the processors in each row or column, and the processor located at one end of the row or column may output data to the external memory, while the other processors may transfer data toward the processor at the one end, and data may be input from the external memory to the processor at the other end. With this configuration, data can be transferred efficiently between processors.
[0021] In the above-described arithmetic device, data may be transferred between a plurality of the external memories. With this configuration, data can be transferred efficiently between the external memories and the processor.
[0022] A data movement method of one aspect of the present disclosure is a data movement method for a computing device having multiple external memories and multiple processors, wherein the processors are associated with the external memories that input and output data, the external memories store data in association with the processors that input and output data, and data is input and output between the processors and the external memories.
[0023] According to the present disclosure, the preparation time for data transfer between an external memory and a processor can be reduced.
[0024] The above and other objects, features, and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. The drawings include: FIG. 1 is a schematic configuration diagram of an accelerator according to an embodiment; FIG. 2 is a schematic configuration diagram showing data transfer between an accelerator external memory, an external memory, and a PE according to an embodiment; FIG. 3 is a schematic configuration diagram showing the layout of an external memory according to an embodiment; FIG. 4 is a schematic configuration diagram showing the layout of an external memory according to an embodiment; FIG. 5 is a schematic configuration diagram showing the layout of an external memory according to an embodiment; FIG. 6 is a schematic configuration diagram showing the swapping of rows and columns of data using an external memory according to an embodiment, where (A) is a diagram before the data swapping and (B) is a diagram after the data swapping; and FIG. 7 is a schematic configuration diagram showing the aggregation of data into a PE using an external memory according to an embodiment, where (A) is a diagram before the data swapping and (B) is a diagram after the data swapping. FIG. 8 is a schematic configuration diagram showing data before aggregation to a PE using an external memory according to an embodiment, where (A) is a diagram showing 8 rows and 8 columns of data stored in an accelerator external memory, and (B) is a diagram showing data held in a PE being output to the external memory. FIG. 9 is a schematic configuration diagram showing data after aggregation to a PE using an external memory according to an embodiment. FIG. 10 is a schematic configuration diagram showing data distribution to a PE using an external memory according to an embodiment, where (A) is a diagram before data distribution and (B) is a diagram after data distribution. FIG. 11 is a schematic configuration diagram showing data distribution to a PE using an external memory according to an embodiment, where (A) is a diagram before data distribution and (B) is a diagram after data distribution. FIG. 12 is a schematic configuration diagram showing data before distribution to a PE using an external memory according to an embodiment. FIG. 13 is a schematic configuration diagram showing data after distribution to a PE using an external memory according to an embodiment. 14A and 14B are schematic diagrams illustrating a case where three-dimensionally arranged data is output from a PE to an external memory according to an embodiment, in which (A) is a diagram illustrating three-dimensionally arranged data stored in an accelerator external memory, and (B) is a diagram illustrating data input from the accelerator external memory to a PE and data output from a PE to the external memory. Fig. 15 is a schematic diagram illustrating data input from an external memory to a PE according to an embodiment.FIG. 16 is a schematic diagram showing three-dimensionally arranged data stored in an accelerator external memory according to an embodiment. FIG. 17 is a schematic diagram showing data output from a PE to an external memory according to an embodiment. FIG. 18 is a schematic diagram showing data input from an external memory to a PE according to an embodiment. FIG. 19 is a schematic diagram showing data output from a PE to an external memory according to an embodiment. FIG. 20 is a schematic diagram showing data input from an external memory to a PE according to an embodiment. FIG. 21 is a schematic diagram showing continuous data movement using two external memories according to an embodiment, where (A) is a diagram before data movement, (B) is a diagram after the first data movement, and (C) is a diagram after the fourth data movement. FIG. 22 is a schematic diagram showing continuous data movement using one external memory according to an embodiment, where (A) is a diagram before data movement, (B) is a diagram after the first data movement, and (C) is a diagram after the second data movement. Fig. 23 is a schematic diagram showing continuous data movement using one external memory according to an embodiment, where (A) is a diagram after the third data movement and (B) is a diagram after the fourth data movement. Fig. 24 is a schematic diagram showing multiple external memories shared by multiple PE groups according to an embodiment. Fig. 25 is a schematic diagram of a conventional arithmetic device.
[0025] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the embodiments described below are examples of how the present disclosure may be implemented, and the present disclosure is not limited to the specific configurations described below. When implementing the present disclosure, specific configurations according to the embodiments may be appropriately adopted.
[0026] FIG. 1 is a schematic diagram of an accelerator 10, which is a computing device according to this embodiment.
[0027] The accelerator 10 is configured with a plurality of PEs 12 and a plurality of external memories 14. In the example of Fig. 1, a total of 16 PEs 12 are arranged two-dimensionally in four rows and four columns, but the accelerator 10 may include any number of PEs 12 as long as the number is more than one. In the following description, the coordinates of the upper left PE 12 are set to (0, 0), and the PEs 12 are also distinguished by their coordinates.
[0028] The PEs 12 constituting the accelerator 10 are electrically connected by wiring 16 for data movement between the PEs 12. A PE 12 can move data between at least other PEs 12 adjacent to it on the top, bottom, left, and right sides via the wiring 16. In this embodiment, data movement is a concept that also includes copying data, and is also expressed as input, output, and transfer of data. Note that although the wiring 16 is omitted in other figures, even in figures where the wiring 16 is omitted, adjacent PEs 12 are still connected by the wiring 16. Furthermore, two or more wirings 16 may be used to connect the PEs 12 to enable simultaneous input and output of multiple data between two PEs 12.
[0029] The PE 12 of this embodiment performs various arithmetic operations such as ==, !=, >, >=, <, <=, >>, <<, or, and, min, max, clip, add, sub, mul, div, mod, macc, etc., using data stored in an internal memory of the PE 12. The numbers written inside the illustrated PE 12 indicate the data stored in the PE 12.
[0030] The external memory 14 is a storage unit that outputs stored data to the PE 12 connected by wiring 16 and stores data input from the PE 12. The PE 12 stores the data input from the external memory 14 in an internal memory provided in the PE 12. The directions of the arrows in Fig. 1 indicate the output direction of data between the external memory 14 and the PE 12 and the output direction of data between the PEs 12.
[0031] Here, the PEs 12 of this embodiment are associated with external memories 14 that input and output data, and the external memories 14 store data in association with the PEs 12 that input and output data. In the example of FIG. 1 , one external memory 14 is associated with the PEs 12 for each row. Each external memory 14 stores data corresponding to the location of the PEs 12 that input and output data. For example, the external memory 14 associated with the PEs 12 in the first row of FIG. 1 stores data such as "00," "01," "02," and "03" from left to right, corresponding to the location of the PEs 12 at coordinates (0,0), (0,1), (0,2), and (0,3). Note that PEs 12 that are not directly connected to the external memory 14 by wiring 16, such as the PEs 12 in the second to fourth columns, receive data via the PEs 12 on the external memory 14 side. For example, the PE 12 at the coordinates (0, 3) inputs and outputs data to and from the external memory 14 via the PEs 12 at the coordinates (0, 2), (0, 1), and (0, 0).
[0032] In the accelerator 10 of this embodiment, the external memory 14 and the PE 12 that transfer data are predetermined, so that the preparation time for data transfer between the external memory 14 and the PE 12, that is, the overhead, can be shortened.
[0033] 2, data is transferred between a memory 20 provided outside the accelerator 10 (hereinafter referred to as an "external accelerator memory") and the external memory 14. Therefore, when data is transferred from the external accelerator memory 20 to the external memory 14, the data is rearranged in accordance with the arrangement positions of the PEs 12 and then stored in the external memory 14. This rearrangement is performed by a rearrangement processing unit 18 provided between the external accelerator memory 20 and the external memory 14. As an example, the rearrangement processing unit 18 is provided for each accelerator 10.
[0034] It is preferable that data movement between the accelerator external memory 20 and the external memory 14 be performed simultaneously with the calculation by the PE 12 .
[0035] 3, the external memory 14 of this embodiment may be arranged in a plurality of different directions relative to the PE 12. This enables data transfer between the PE 12 and the external memory 14 in various directions, allowing more data to be transferred at one time.
[0036] 3, four pieces of data can be allocated to each of the internal memories of a plurality of PEs 12 arranged in two rows and two columns. Then, external memories 14 are arranged in four directions relative to each PE 12. By arranging the external memories 14 in this manner, the accelerator 10 can input data from the external memories 14 to the PEs 12 simultaneously from four directions, and output data from the PEs 12 to the external memories 14 simultaneously in four directions.
[0037] In the example of FIG. 3 , external memories 14 are arranged to the left and above the PE 12 at coordinates (0,0). The external memory 14 to the left of the PE 12 at coordinates (0,0) outputs “00,” one of the data associated with the PE 12 at coordinates (0,0), and “01,” one of the data associated with the PE 12 at coordinates (0,1). The external memory 14 above the PE 12 at coordinates (0,0) outputs “20,” one of the data associated with the PE 12 at coordinates (0,0), and “30,” one of the data associated with the PE 12 at coordinates (1,0). That is, the external memories 14 arranged to the left and right of the PE 12 output data to the PE 12 in the same row. The external memories 14 arranged above and below the PE 12 output data to the PE 12 in the same column. The external memories 14 may be arranged in a diagonal direction, or in a multidimensional direction of two or more dimensions, such as three-dimensional, four-dimensional, or five-dimensional, relative to the PE 12.
[0038] 4, the external memories 14 may be arranged in a one-to-one relationship with the PEs 12. For this purpose, the external memories 14 and the PEs 12 are connected in a one-to-one relationship by wiring 16. In this case, it is preferable to arrange the external memories 14 near the PEs 12 so that the wiring 16 does not become too long. The number of external memories 14 may be less than the number of rows of the PEs 12. For example, if the number of rows of the PEs 12 is four, the number of external memories 14 may be two, which is half the number of rows of the PEs 12, and one external memory 14 may be shared by two rows of the PEs 12.
[0039] 5, the plurality of PEs 12 may be divided into a plurality of PE groups 30 in a row or a column, and the external memory 14 may be arranged for each PE group 30. In the example of FIG. 5, one row consisting of four PEs 12 is divided into two PE groups 30, and an external memory 14 is provided for each PE group 30. The external memory 14 may be divided into PE groups 30 in a plurality of rows or a plurality of columns, or both, and the external memory 14 may be arranged for each PE group 30. That is, the plurality of PEs 12 may be divided into a plurality of PE groups 30 in at least one of a plurality of rows and a plurality of columns, and the external memory 14 may be arranged for each PE group 30. For example, PEs 12 arranged in two rows and two columns may be grouped into one PE group 30, and PEs 12 arranged in four rows and four columns may be grouped into PE groups 30 in two rows and two columns, and the external memory 14 may be arranged for four PE groups 30.
[0040] (Data Swapping in Row and Column Direction) The accelerator 10 of this embodiment may use the external memory 14 to swap rows and columns of data held in the PEs 12. That is, data transferred to the external memory 14 from multiple PEs 12 arranged in the same row may be input from the external memory 14 to multiple PEs 12 arranged in the same column. Furthermore, data transferred to the external memory 14 from multiple PEs 12 arranged in the same column may be input from the external memory 14 to multiple PEs 12 arranged in the same row. This makes it possible to easily swap rows and columns of data held in the PEs 12.
[0041] 6 shows an example of shuffling rows and columns of data using the external memory 14. Fig. 6(A) shows the state before the data shuffling, in which data held in the PE 12 is output row by row to the external memory 14. Fig. 6(B) shows the state after the data shuffling, in which data stored row by row from the external memory 14 is input column by column to the PE 12. In the example of Fig. 6, data from the PE 12 in the first row moves to the PE 12 in the first column, data from the PE 12 in the second row moves to the PE 12 in the second column, data from the PE 12 in the third row is output to the PE 12 in the third column, and data from the PE 12 in the fourth row moves to the PE 12 in the fourth column.
[0042] When data in the same column is swapped to the same row, multiple external memories 14 are provided for each column in association with the PE 12. Data for each column of the PE 12 is output to the external memories 14, and data for each row of the PE 12 is input from the external memories 14.
[0043] (Data Aggregation) Data output from the plurality of PEs 12 to the external memory 14 in this embodiment may be input from the external memory 14 to the PEs 12 so that the PEs 12 hold the plurality of data. This allows the accelerator 10 in this embodiment to aggregate the plurality of data in the PEs 12 and hold them.
[0044] 7 is a schematic configuration diagram showing an example of data aggregation to the PEs 12 using the external memory 14. FIG. 7A shows data before aggregation, with data being output row by row from the PEs 12 to the external memory 14. FIG. 7B shows data after aggregation, with the accelerator 10 inputting data row by row from the external memory 14 to the PEs 12. In the example of FIG. 7, the accelerator 10 inputs two pieces of data from the external memory 14 for each PE 12, thereby aggregating the data in the PEs 12.
[0045] Specifically, data "00" to "03" are output from the PEs 12 at coordinates (0,0) to (0,3) to the external memory 14 corresponding to the PE 12 in the first row. The accelerator 10 then inputs data "00" and "01" from the external memory 14 to the PE 12 at coordinates (0,0), and inputs data "02" and "03" to the PE 12 at coordinates (0,1). Data is similarly aggregated for the PEs 12 in the other rows. In this way, in the example of FIG. 7 , the data is aggregated in the PEs 12 by causing the PEs 12 corresponding to the external memory 14 to which data is output from the multiple PEs 12 to store multiple data. Note that in the example of FIG. 7 , the number of PEs 12 that store data is reduced after the data aggregation compared to before the aggregation.
[0046] 7, data from two PEs 12 is aggregated into one PE 12, but the data from more PEs 12, such as three, four, eight, or all of the data in the same row, may be aggregated into a fewer number of PEs 12. Furthermore, when aggregating, the data may be simultaneously converted into average, maximum, or minimum data, or may be reduced to one of the data.
[0047] (Data Aggregation 2) Next, another example of data aggregation using the external memory 14 will be described with reference to Figures 8 and 9. Figure 8 is a schematic diagram showing the data before aggregation into the PE 12 using the external memory 14, and Figure 9 is a schematic diagram showing the data after aggregation.
[0048] 8A shows data of 8 rows and 8 columns stored in the accelerator external memory 20. The 8 rows and 8 columns of data are divided into groups 34A to 34B of 4 rows and 4 columns of data. In this way, FIGS. 8 and 9 illustrate an example where data of 8 rows and 8 columns is arranged in PEs 12 of 4 rows and 4 columns, that is, where more data than the number of PEs 12 is arranged in the PEs 12.
[0049] The accelerator 10 inputs data of 8 rows and 8 columns stored in the accelerator external memory 20 one by one, starting with the PE 12 at the upper left coordinate (0,0). Therefore, data of the same row is input to the right, and data of the same column is input to the bottom. Specifically, after the accelerator 10 inputs data of the same row from the left end to the right end PE 12, it again inputs data from the left end PE 12 as a second round. Similarly, after the accelerator 10 inputs data of the same column from the top end to the bottom end PE 12, it again inputs data from the top end PE 12 as a second round.
[0050] In this method of allocating data one by one to the PE 12 arranged in four rows and four columns, for example, data "00" in FIG. 8A is input as the first round of data to the PE 12 at coordinates (0,0). Then, data "04" to the right of data "03" in FIG. 8A becomes the second round of data in the same row and is input to the PE 12 at coordinates (0,0) on the left side. Also, data "40" below data "30" in FIG. 8A becomes the second round of data in the same column and is input to the PE 12 at coordinates (0,0). Furthermore, data "44" to the right of data "43" in FIG. 8A becomes the second round of data in the same matrix and is input to the PE 12 at coordinates (0,0).
[0051] In this embodiment, data held at the same address (hereinafter referred to as a "relative address") in the internal memories of multiple PEs 12 is considered to be included in the same data group (hereinafter referred to as a "tile"). In this embodiment, in the internal memory of each PE 12, the relative address indicated in the first row and first column is considered to be tile 13A, the relative address indicated in the first row and second column is considered to be tile 13B, the relative address indicated in the second row and first column is considered to be tile 13C, and the relative address indicated in the second row and second column is considered to be tile 13D.
[0052] The data of group 34A in the accelerator external memory 20 is input to the internal memory of each PE 12 as tile 13A. The data of group 34B in the accelerator external memory 20 is input to the internal memory of each PE 12 as tile 13B. The data of group 34C in the accelerator external memory 20 is input to the internal memory of each PE 12 as tile 13C. The data of group 34D in the accelerator external memory 20 is input to the internal memory of each PE 12 as tile 13D.
[0053] In this way, the data in 8 rows and 8 columns stored in the accelerator external memory 20 is allocated to the PEs 12 in 4 rows and 4 columns, so that each PE 12 has four pieces of data allocated to it.
[0054] 8B, the data in two rows and two columns held by each of the 4-row, 4-column PEs 12, i.e., the data in eight rows and eight columns, is output row by row to the external memories 14A to 14H. More specifically, the PE 12 in the first row outputs the data of tiles 13A and 13B to the external memory 14A, and outputs the data of tiles 13C and 13D to the external memory 14B. At this time, the PE 12 in the first row outputs the data of tile 13A to the external memory 14A, and then outputs the data of tile 13B to the external memory 14A, and then outputs the data of tile 13C to the external memory 14B, and then outputs the data of tile 13D to the external memory 14B.
[0055] Similarly, the PE 12 in the second row outputs the data of tiles 13A and 13B to external memory 14C, and outputs the data of tiles 13C and 13D to external memory 14D. The PE 12 in the third row outputs the data of tiles 13A and 13B to external memory 14E, and outputs the data of tiles 13C and 13D to external memory 14F. The PE 12 in the fourth row outputs the data of tiles 13A and 13B to external memory 14G, and outputs the data of tiles 13C and 13D to external memory 14H.
[0056] Next, as shown in FIG. 9 , the accelerator 10 aggregates data by inputting a total of four data items, two rows and two columns, from the external memory 14 to each PE 12. In the example of FIG. 9 , the accelerator 10 inputs data stored in the external memories 14A and 14C to the PE 12 in the first row. More specifically, the accelerator 10 inputs data “00” and “01” from the external memory 14A and data “10” and “11” from the external memory 14C to the PE 12 at coordinates (0,0). The accelerator 10 also inputs data “02” and “03” from the external memory 14A and data “12” and “13” from the external memory 14C to the PE 12 at coordinates (0,1). The accelerator 10 also inputs data “04” and “05” from the external memory 14A and data “14” and “15” from the external memory 14C to the PE 12 at coordinates (0,2). The accelerator 10 also inputs the data "06" and "07" from the external memory 14A and the data "16" and "17" from the external memory 14C to the PE 12 at the coordinates (0, 3).
[0057] Similarly, the accelerator 10 inputs four pieces of data stored in the external memories 14E and 14G to the PEs 12 in the second row. The accelerator 10 also inputs four pieces of data stored in the external memories 14B and 14D to the PEs 12 in the third row. The accelerator 10 also inputs four pieces of data stored in the external memories 14F and 14H to the PEs 12 in the fourth row.
[0058] That is, the accelerator 10 aggregates data of 2 rows and 2 columns into the PE 12 of 4 rows and 4 columns so as to reproduce the arrangement of data of 8 rows and 8 columns in the accelerator external memory 20 .
[0059] 8 and 9, the external memory 14 stores the data in the order in which it is output from the PE 12, and the data is rearranged when it is output from the external memory 14 to the PE 12, but this is not limiting, and the data may be rearranged when it is output from the PE 12 to the external memory 14. Also, the data may be partially rearranged when it is output from the PE 12 to the external memory 14, and the rest of the data may be rearranged when it is output from the external memory 14 to the PE 12. For example, the top and bottom of the data may be rearranged when it is output from the PE 12 to the external memory 14, and the left and right of the data may be rearranged when it is output from the external memory 14 to the PE 12.
[0060] Furthermore, when rearranging data in reverse order, this may be done simultaneously when outputting data from the PE 12 to the external memory 14 or when inputting data from the external memory 14 to the PE 12. In this embodiment, the data is aggregated into two areas each for rows and columns of the internal memory provided in the PE 12, but the number of areas is not limited to two, and data may be aggregated into three, eight, or more areas. Furthermore, when aggregating, the data may be simultaneously converted into average, maximum, or minimum data, or may be reduced to any one of the data.
[0061] 8 and 9, data transfer between PE 12 and external memory 14 is described for each row, but data transfer between PE 12 and external memory 14 may also be performed for each column, or data transfer between PE 12 and external memory 14 may also be performed for both rows and columns simultaneously.
[0062] 8 and 9 are described using eight external memories 14, which is twice the number of rows in the PE 12 (four), but the number of external memories 14 may be four, the same as the number of rows in the PE 12, and data may be transferred at different timings. In this case, for example, external memories 14A and 14B may be combined into one external memory 14. Similarly, external memories 14C and 14D may be combined into one external memory 14, external memories 14E and 14F may be combined into one external memory 14, and external memories 14G and 14H may be combined into one external memory 14. Also, the number of external memories 14 may be less than the number of rows in the PE 12.
[0063] (Data Distribution) In this embodiment, data output from the multiple PEs 12 to the external memory 14 is copied and input from the external memory 14 to the PEs 12 so that the multiple PEs 12 hold the same data. This allows the accelerator 10 in this embodiment to distribute multiple pieces of data to the PEs 12 and have them hold the data.
[0064] 10 shows an example in which one piece of data is copied and distributed to the PEs 12 in one row and two columns. Fig. 10(A) shows the data before distribution, and the data held by the PEs 12 in four rows and four columns is output row by row to the external memory 14. Then, as shown in Fig. 10(B), the accelerator 10 copies and inputs the data from the external memory 14 to each of the PEs 12 in two columns.
[0065] 10, data "00" to "03" are output from the PEs 12 at coordinates (0,0) to (0,3) to the external memory 14A corresponding to the PEs 12 in the first row. Then, data "00" is input from the external memory 14A to the PEs 12 at coordinates (0,0) and (0,1), and data "01" is input to the PEs 12 at coordinates (0,2) and (0,3). Data is similarly distributed to the PEs 12 in the other rows.
[0066] Next, it is assumed that data "02" is copied and distributed to the PEs 12 at coordinates (0,4) and (0,5). However, in FIG. 10, four PEs 12 are arranged in the same row, so distribution is not possible. In this case, the accelerator 10 allocates data starting from the PE 12 at (0,0) in the second round. That is, the accelerator 10 distributes data "02" to the PEs 12 at coordinates (0,0) and (0,1), and distributes data "03" to the PEs 12 at coordinates (0,2) and (0,3). Therefore, each PE 12 holds a total of two pieces of data in tiles 13A and 13B of the internal memory. Similarly, data held in the external memory 14 is input to each PE 12.
[0067] 10, data aggregation and distribution are performed for the same row, but the order of columns for the same row may be reversed, or data aggregation and distribution may be performed in reverse order. Also, in the example of Fig. 10, one piece of data is distributed to two PEs 12, but data may be distributed to more PEs 12, such as three, four, eight, or all of the PEs in the same row.
[0068] (Data Distribution 2) Fig. 11 shows an example in which one piece of data is copied and distributed to two rows and two columns of PEs 12, with Fig. 11(A) showing the data before distribution and Fig. 11(B) showing the data after distribution. In the example of Fig. 11, data held by the PEs 12 is output row by row to the external memory 14. Then, one piece of data is copied from the external memory 14 and input to each of the two rows and two columns of PEs 12.
[0069] 11B, the accelerator 10 inputs the data "00" output from the PE 12 at the coordinate (0,0) to the external memory 14A to the PEs 12 at the coordinates (0,0), (0,1), (1,0), and (1,1). The accelerator 10 inputs the data "01" output from the PE 12 at the coordinate (0,1) to the external memory 14A to the PEs 12 at the coordinates (0,2), (0,3), (1,2), and (1,3).
[0070] Note that it is desired to input the data "02" output from the PE 12 at coordinates (0,2) to the external memory 14A to the PEs 12 at coordinates (0,4), (0,5), (1,4), and (1,5). However, since there are only four columns of PEs 12, these would be PEs 12 at coordinates outside the range. Therefore, if the data output destination is a PE 12 outside the range, the accelerator 10 inputs data again from the column with the smallest coordinates as a second round. In other words, the accelerator 10 inputs the data "02" to the PEs 12 at coordinates (0,0), (0,1), (1,0), and (1,1). Similarly, the accelerator 10 inputs the data "03" output from the PE 12 at coordinates (0,3) to the external memory 14A to the PEs 12 at coordinates (0,2), (0,3), (1,2), and (1,3). Similarly, the data held in the external memories 14B to 14D are input to each PE 12. As a result, each PE 12 holds a total of four pieces of data in the tiles 13A to 13D of the internal memory.
[0071] 12 and 13, which will be described next, are also examples in which one piece of data is copied and distributed to two rows and two columns of PEs 12. In the examples of Figures 12 and 13, the PEs 12 in each row output data to two external memories 14, and twice as many columns of data, i.e., four pieces of data, are input to the PEs 12 from each of the two external memories 14. Figure 12 shows the data before distribution, with the PEs 12 in the first row outputting data to external memories 14A and 14C, the PEs 12 in the second row outputting data to external memories 14E and 14G, the PEs 12 in the third row outputting data to external memories 14B and 14D, and the PEs 12 in the fourth row outputting data to external memories 14F and 14H.
[0072] 13, the accelerator 10 inputs data of two rows and two columns from two external memories 14 to each of the PEs 12 in one associated row. In the example of FIG. 13, the accelerator 10 inputs data "00", "00", "01", and "01" from the external memory 14A to the PEs 12 at coordinates (0,0) to (0,3). Note that if the PEs 12 have more columns, the data "02" to "03" would be input to the PEs 12 at coordinates (0,4) to (0,7). However, in the example of FIG. 13, the number of columns of the PEs 12 is four, which falls outside the range, so the accelerator 10 again inputs data from the column with the smaller coordinates in the second round. In other words, the accelerator 10 inputs data "02", "02", "03", and "03" from the external memory 14A to the PEs 12 at coordinates (0,0) to (0,3). Furthermore, the accelerator 10 inputs data "20", "20", "21", "21" and data "22", "22", "23", "23" from the external memory 14B to the PEs 12 at coordinates (0,0) to (0,3).
[0073] Similarly, the accelerator 10 inputs four pieces of data for each PE 12 in the second row from the external memories 14C and 14D, inputs four pieces of data for each PE 12 in the third row from the external memories 14E and 14F, and inputs four pieces of data for each PE 12 in the fourth row from the external memories 14G and 14H.
[0074] Furthermore, when the accelerator 10 wants to rearrange the data in reverse order, it may do so at the same time as inputting the data from the external memory 14 to the PE 12. In this embodiment, the data is distributed to two areas each for rows and columns of the internal memory provided in the PE 12, but the data may be distributed to more areas, such as three, four, or eight, without being limited to two.
[0075] 12 and 13, eight external memories 14 are used, which is twice the number of rows in the PE 12 (four). However, the number of external memories 14 may be four, the same as the number of rows in the PE 12, and data may be transferred at different timings. In this case, for example, 14A and 14B may be combined into one external memory. Similarly, 14C and 14D, 14E and 14F, and 14G and 14H may be combined into one external memory. The number of external memories 14 may also be less than the number of rows in the PE 12.
[0076] (Input 1 of Three-Dimensional Arrayed Data to Multiple PEs) In the accelerator 10 of this embodiment, data that has been made into a multidimensional array by multiple groups (hereinafter referred to as "channels") 21 formed of two-dimensionally arrayed data is stored in the accelerator external memory 20. The accelerator 10 inputs the data in the nth row and mth column of the multiple channels 21 from the accelerator external memory 20 to the PE 12 in the nth row and mth column, and then causes the PE 12 to output the data to the external memory 14 corresponding to the PE 12. Then, the external memory 14 inputs the data to the multiple PEs 12 so that the arrangement of the data that has been made into a multidimensional array by the multiple channels 21 becomes a two-dimensional array.
[0077] An example of the channels 21 will now be described with reference to FIG. 14. In the example of FIG. 14, four channels 21-0 to 21-3 are stored in the accelerator external memory 20, and each of the channels 21-0 to 21-3 contains four pieces of data arranged in two rows and two columns. In other words, the data arrangement in each channel 21 is two-dimensional, and the channels 21 0 to 3 represent the third dimension. In this way, the third dimension is called a "channel," and the number of the third dimension is called the "number of channels." In the example of FIG. 14, the number of channels is four.
[0078] The data in the two matrices of channels 21-0 to 21-3 are then input to the PE 12 in two rows and two columns. FIG. 14B shows the arrangement of the data input to the PE 12 and the data stored in the external memory 14 from the PE 12. In the example of FIG. 14B, the data "000," "100," "200," and "300" in the first row and first column of channels 21-0 to 21-3 are input to tiles 13A to 13D of the internal memory of the PE 12 at coordinates (0,0) in the first row and first column. Furthermore, the data "001," "101," "201," and "301" in the first row and second column of channels 21-0 to 21-3 are input to tiles 13A to 13D of the internal memory of the PE 12 at coordinates (0,1) in the first row and second column. Furthermore, the data "010", "110", "210", and "310" in the second row and first column of channels 21-0 to 21-3 are input to tiles 13A to 13D of the internal memory of PE 12 at coordinates (1,0) in the second row and first column. Furthermore, the data "011", "111", "211", and "311" in the second row and second column of channels 21-0 to 21-3 are input to tiles 13A to 13D of the internal memory of PE 12 at coordinates (1,1) in the second row and second column.
[0079] Then, the PE 12 outputs a plurality of data to the corresponding external memory 14. In the example of Fig. 14B, among the data of the PE 12 in the first row, the data of tiles 13A and 13B is output to the external memory 14A, and the data of tiles 13C and 13D is output to the external memory 14C. Furthermore, among the data of the PE 12 in the second row, the data of tiles 13A and 13B is output to the external memory 14B, and the data of tiles 13C and 13D is output to the external memory 14D.
[0080] 15 is a diagram showing a case where data output to the external memory 14 in FIG. 14 is input to the PE 12. As described above, the external memory 14 inputs data to the multiple PEs 12 so that the data arrangement, which was a three-dimensional array by the channels 21-0 to 21-3, becomes a two-dimensional array. That is, in the example of FIG. 15, data is input from the external memory 14A to the PE 12 in the first row, data is input from the external memory 14B to the PE 12 in the second row, data is input from the external memory 14C to the PE 12 in the third row, and data is input from the external memory 14D to the PE 12 in the fourth row. Each PE 12 stores the input data in a tile 13A.
[0081] As a result, the four PEs 12 in the first row hold the data "000", "001", "100", and "101", the four PEs 12 in the second row hold the data "010", "011", "110", and "111", the four PEs 12 in the third row hold the data "200", "201", "300", and "301", and the four PEs 12 in the fourth row hold the data "210", "211", "310", and "311". Such a two-dimensional array of data is a two-dimensional array of the three-dimensional array of data represented by channels 21-0 to 21-3.
[0082] In the example of FIG. 14 , the number of rows of the PE 12 is four, and the number of rows of data stored in each PE 12 from the accelerator external memory 20 is two. Therefore, the transfer destinations of the data stored in the tiles 13A to 13D of the PE 12 in the first row of FIG. 14 are the PEs 12 in the first and third rows of FIG. 15 . That is, the data transfer destination of the PE 12 in the first row is the number of rows indicated by the first row of the PE 12 and the number of rows (n / 2+1) of the PE 12. Similarly, the data transfer destination of the tiles 13A to 13D of the PE 12 in the second row is the number of rows indicated by the second row of the PE 12 and the number of rows (n / 2+2) of the PE 12. The location of the PE 12 to which the data of the PE 12 is transferred via the external memory 14 may be determined by other methods as long as the arrangement of the data, which has been arranged in a three-dimensional array by the multiple channels 21, can be arranged in a two-dimensional array.
[0083] 14 and 15 show an example in which data of two rows and two columns held by four channels 21 is aggregated into one tile 13A held by a PE 12 with four rows and four columns, but the number of PEs 12 to aggregate data and the number of tiles 13 are not limited to this. For example, in a PE 12 with 32 rows and 32 columns, data held in 64 tiles 13 with four rows and four columns may be aggregated into one tile 13, or data held in 16 tiles 13 with eight rows and eight columns may be aggregated into one tile 13.
[0084] In addition, when the number of rows of data is eight in 32-by-32 PEs 12, the data transfer destination of the PE 12 in the first row of each tile 13 may be two rows: the first row of the PE 12 and the 17th row (n / 2+1 = the number of rows of the PE 12). Alternatively, the data transfer destination may be four rows: the first row of the PE 12, the 9th row (n / 4+1 = the number of rows of the PE 12), the 17th row (n / 4×2+1 = the number of rows of the PE 12), and the 25th row (n / 4×3+1 = the number of rows of the PE 12). In the case of an even fewer number of rows of data (four rows), the data transfer destination may be eight rows (n / 8×(0 to 7)+1 = the number of rows of the PE 12).
[0085] 14 and 15, the number of pieces of data three-dimensionally arranged by the multiple channels 21 is the same as the number of PEs 12. Next, with reference to FIGS. 16 to 18 and 19 and 20, a case will be described in which the number of pieces of data three-dimensionally arranged by the multiple channels 21 is greater than the number of PEs 12 included in the accelerator 10.
[0086] Fig. 16 shows three-dimensionally arranged data stored in the accelerator external memory 20. In the example of Fig. 16, data arranged two-dimensionally in four rows and four columns is defined as one channel 21, and data is arranged three-dimensionally by four channels 21-0 to 21-3. That is, 64 pieces of data are arranged three-dimensionally.
[0087] 17 is a diagram showing data output from the PE 12 to the external memory 14. In the example of FIG. 17, data is held in all of the PEs 12 arranged in four rows and four columns. Therefore, the tile 13A of the internal memory of the PE 12 arranged in the nth row and mth column holds the data of the nth row and mth column of channel 21-0. Furthermore, the tile 13B of the PE 12 arranged in the nth row and mth column holds the data of the nth row and mth column of channel 21-1. The tile 13C of the PE 12 arranged in the nth row and mth column holds the data of the nth row and mth column of channel 21-2. The tile 13D of the PE 12 arranged in the nth row and mth column holds the data of the nth row and mth column of channel 21-3.
[0088] Then, data is output to the external memory 14 corresponding to each column of PEs 12. In the example of FIG. 17 , data from PEs 12 in the first column is output to external memory 14A, data from PEs 12 in the second column is output to external memory 14B, data from PEs 12 in the third column is output to external memory 14C, and data from PEs 12 in the fourth column is output to external memory 14D. At this time, the data from each PE 12 is output to each row of the external memory 14 in the order of tiles 13A to 13D. That is, data from PEs 12 in the first row is output to the first row of the external memory 14, data from PEs 12 in the second row is output to the second row of the external memory 14, data from PEs 12 in the third row is output to the third row of the external memory 14, and data from PEs 12 in the fourth row is output to the fourth row of the external memory 14.
[0089] 18 is a diagram showing the state after the data has been rearranged and input from the external memories 14 shown in FIG. 17 to the PEs 12. Each external memory 14 inputs data to the PEs 12 in the corresponding column. That is, external memory 14A inputs data to the PEs 12 in the first column, external memory 14B inputs data to the PEs 12 in the second column, external memory 14C inputs data to the PEs 12 in the third column, and external memory 14D inputs data to the PEs 12 in the fourth column.
[0090] The data in the first column of the external memory 14 is input to the PE 12 in the first row, the data in the second column of the external memory 14 is input to the PE 12 in the second row, the data in the third column of the external memory 14 is input to the PE 12 in the third row, and the data in the fourth column of the external memory 14 is input to the PE 12 in the fourth row. Therefore, the tiles 13A to 13D of each PE 12 are arranged in one column, and the data for each column of the external memory 14 is input to these tiles 13A to 13D.
[0091] As a result, the data arrangement by the PE 12 is such that the three-dimensional array of data represented by channels 21-0 to 21-3 is converted into a two-dimensional array that is different from the data before output and before input. That is, the data of channel 21-0 shown in Fig. 16 is arranged in the PE 12 in the first row shown in Fig. 18, the data of channel 21-1 is arranged in the PE 12 in the second row, the data of channel 21-2 is arranged in the PE 12 in the third row, and the data of channel 21-3 is arranged in the PE 12 in the fourth row. In this way, in the example of Fig. 18, the data of each channel 21 is arranged in each row of the PE 12.
[0092] Next, another embodiment will be described with reference to FIGS.
[0093] In the example of Fig. 19, the data of the three-dimensional array shown in Fig. 16 is input to the PEs 12 arranged in four rows and four columns, as in Fig. 17, and data is output to the external memory 14 corresponding to each row of the PEs 12. In the example of Fig. 19, data from the PEs 12 in the first row is output to the external memory 14A, data from the PEs 12 in the second row is output to the external memory 14B, data from the PEs 12 in the third row is output to the external memory 14C, and data from the PEs 12 in the fourth row is output to the external memory 14D. At this time, the data from each PE 12 is output to each column of the external memory 14 in the order of tiles 13A to 13D.
[0094] 20 is a diagram showing the state after the data shown in FIG. 19 has been rearranged and input to the PEs 12. The data output to the external memories 14A to 14D is input to the PEs 12 in the corresponding rows from each external memory 14. That is, the external memory 14A inputs data to the PEs 12 in the first row, the external memory 14B inputs data to the PEs 12 in the second row, the external memory 14C inputs data to the PEs 12 in the third row, and the external memory 14D inputs data to the PEs 12 in the fourth row.
[0095] The data in the first row of the external memory 14 is input to the PE 12 in the first column, the data in the second row of the external memory 14 is input to the PE 12 in the second column, the data in the third row of the external memory 14 is input to the PE 12 in the third column, and the data in the fourth row of the external memory 14 is input to the PE 12 in the fourth column. Therefore, the tiles 13A to 13D of each PE 12 are one row, and the data for each row of the external memory 14 is input to these tiles 13A to 13D.
[0096] As a result, the data arrangement by the PE 12 is such that the three-dimensional array of data represented by channels 21-0 to 21-3 is converted into a two-dimensional array that is different from the data before output and before input. That is, the data of channel 21-0 shown in Fig. 16 is arranged in the PE 12 in the first column shown in Fig. 20, the data of channel 21-1 is arranged in the PE 12 in the second column, the data of channel 21-2 is arranged in the PE 12 in the third column, and the data of channel 21-3 is arranged in the PE 12 in the fourth column. In this way, in the example of Fig. 20, the data of each channel 21 is arranged for each column of PE 12.
[0097] Above, with reference to Figures 14 to 20, we have explained a form in which data arranged three-dimensionally by multiple channels 21 is rearranged via external memory 14 and input to PE 12 so that it becomes a two-dimensional array that is different from the array before output and before input. However, the method of assigning the three-dimensional arrangement of data channels 21, rows, and columns to the two-dimensional arrangement of rows and columns of PE 12 is not limited to the above method, and rearrangement may be performed by other methods.
[0098] Furthermore, the data stored in the accelerator external memory 20 may be multidimensional (four or more dimensions) instead of three-dimensional. For example, a plurality of groups each having a plurality of channels 21 may be further set, and these groups may be the fourth or higher dimension. Conversely, the arrangement of the PEs 12 may be three or more dimensions, and two-dimensional array data may be input to the PEs 12 with a three or higher dimensional array for rearrangement.
[0099] (Continuous Movement of Data Between External Memory and PEs) The accelerator 10 of this embodiment associates one or more external memories 14 with the PEs 12 of each row or column, and the PEs 12 located at one end of the row or column output data to the external memory 14, while the other PEs 12 move data toward the PE 12 at the end, and data is input to the PE 12 at the other end from the external memory 14. In other words, the PEs 12 of each row or column move data between themselves and the external memory 14 in turn, and therefore the accelerator 10 of this embodiment can efficiently move data between the PEs 12.
[0100] In the example of Figure 21, the PE 12 in the first row is associated with external memories 14A1 and 14A2, the PE 12 in the second row is associated with external memories 14B1 and 14B2, the PE 12 in the third row is associated with external memories 14C1 and 14C2, and the PE 12 in the fourth row is associated with external memories 14D1 and 14D2.
[0101] The external memories 14A1 to 14D1 are arranged on the left side of each row, and the external memories 14A2 to 14D2 are arranged on the right side of each row. The PE 12 in the first column, which is the leftmost column in each matrix, receives data from the external memories 14A1 to 14D1 on the left side. The PE 12 in the fourth column, which is the rightmost column in each matrix, outputs data to the external memories 14A2 to 14D2 on the right side.
[0102] 21A shows the state before the data movement, (B) shows the state after the first data movement, and (C) shows the state after the fourth data movement. In the first data movement, the PE 12 in the first column, which is the PE 12 at one end of each row, receives data "03," "13," "23," and "33" from the external memories 14A1 to 14D1 on the left side, starting from the top. At the same time, data held in the other PEs 12 is moved to the PE 12 on the right side, and the PE 12 in the fourth column, which is the PE 12 at the other end, outputs data "07," "17," "27," and "37" to the right side external memories 14A2 to 14D2.
[0103] By repeating this process, as shown in FIG. 21(C), all of the data stored in external memories 14A1 to 14D1 is moved to PE 12, and the data held in PE 12 is stored in external memories 14A2 to 14D2.
[0104] In the above example, two external memories 14 are associated with each row or column, but this is not limited to this. One external memory 14 may be associated with each row or column, and data may be moved using one external memory 14.
[0105] 22 and 23 are schematic diagrams showing continuous data transfer using one external memory 14 for each row. Fig. 22(A) shows the state before data transfer, Fig. 22(B) shows the state after the first data transfer, Fig. 22(C) shows the state after the second data transfer, Fig. 23(A) shows the state after the third data transfer, and Fig. 23(B) shows the state after the fourth data transfer.
[0106] In the first data movement, the PE 12 in the first column, which is the PE 12 at one end of each row, receives data "03", "13", "23", and "33" from the external memories 14A to 14D on the left side. At the same time, data held in the other PEs 12 moves to the PE 12 on the right side, and the PE 12 in the fourth column, which is the PE 12 at the other end, outputs data "07", "17", "27", and "37" to the external memories 14A to 14D. At this time, the data stored in the external memories 14A to 14D also moves rightward, i.e., toward the PE 12, and the data output from the PE 12 in the fourth column is stored at the left end farthest from the PE 12.
[0107] By repeating this continuous movement four times, as shown in FIG. 23B, the data held in the PE 12 is swapped with the data stored in the external memories 14A to 14D. Also, as shown in FIGS. 22 and 23, the data to be moved to the PE 12 is stored in the external memory 14 on the side closer to the PE 12, and the data output from the PE 12 is stored on the side farther from the PE 12. This shortens the time required to input data to the PE 12. In this embodiment, in order to efficiently use the external memory 14, the PE 12 outputs data to an area that becomes vacant after inputting data to the PE 12. However, this is not limiting. There are also cases where data is desired to be left in the external memory 14, such as when reusing data input to the PE 12. Therefore, the external memory 14 may retain the data input to the PE 12, and the PE 12 may output the data to a different area of the external memory 14.
[0108] While the present disclosure has been described above using the above-described embodiments, the technical scope of the present disclosure is not limited to the scope described in the above-described embodiments. Various modifications or improvements can be made to the above-described embodiments without departing from the gist of the disclosure, and such modifications or improvements are also included in the technical scope of the present disclosure.
[0109] For example, in the above embodiment, the direction of data movement between the external memory 14 and the PE 12 may be changed from left to right, data may be moved not only in the same row direction but also in the same column direction, or data may be moved in the same row direction and the same column direction simultaneously.
[0110] The accelerator 10 of this embodiment may transfer data between a plurality of external memories 14. For example, data may be transferred from one external memory 14 to another external memory 14 depending on the content of data movement between the external memory 14 and the PE 12. This allows the accelerator 10 of this embodiment to efficiently transfer data between the external memory 14 and the PE 12.
[0111] Furthermore, data may be moved from the external memory 14 to the PE 12 simultaneously with the calculation performed by the PE 12. In this case, the PE 12 may have a buffer for data transfer.
[0112] Furthermore, if the distance between the external memory 14 and the PE 12 is so great that data cannot be moved in one processing run, a buffer may be provided between the external memory 14 and the PE 12, and data may be moved by a pipeline.
[0113] Furthermore, the amount of data that can be moved at one time may be increased by multiplexing the wiring 16 between the external memory 14 and the PEs 12 and the wiring 16 between the PEs 12 .
[0114] 24, the PE group 30 may be one system, and the external memory 14 may be shared by a plurality of systems. With such a configuration, for example, data may be transferred between systems via the external memory 14.
[0115] Furthermore, the data arrangement in one external memory 14 and the data arrangement in one PE 12 shown in each figure are examples, and are arranged as shown in the figure for ease of explanation, but they do not necessarily have to be in two dimensions or in the order shown in the figure, and it is sufficient to arrange the data in a way that makes it easy to output and input.
[0116] Next, the features of the present disclosure are as follows.
[0117] (Aspect 1) A computing device (10) comprising a plurality of external memories (14) and a plurality of processors (12), wherein the processors are associated with the external memories that input and output data, and the external memories store data in association with the processors that input and output data.
[0118] (Aspect 2) The arithmetic device according to aspect 1, wherein data transferred from an external memory (20) provided outside the arithmetic device to the external memory is sorted in accordance with the location of the processor and stored in the external memory.
[0119] (Aspect 3) The arithmetic device according to aspect 1 or aspect 2, wherein the external memories are arranged in a plurality of different directions with respect to the processor.
[0120] (Aspect 4) The arithmetic device according to aspect 1 or aspect 2, wherein the external memory is arranged in a one-to-one correspondence with the processor.
[0121] (Aspect 5) The arithmetic device according to aspect 1 or aspect 2, wherein the plurality of processors are divided into a plurality of groups (30) in one row or one column, and the external memory is arranged for each of the groups.
[0122] (Aspect 6) The arithmetic device according to aspect 1 or aspect 2, wherein the plurality of processors are divided into a plurality of groups in at least one of a plurality of rows and a plurality of columns, and the external memory is arranged for each of the groups.
[0123] (Aspect 7) The arithmetic device according to any one of Aspects 1 to 6, wherein data output to the external memory from a plurality of the processors arranged in the same row is input from the external memory to a plurality of the processors arranged in the same column, and data output to the external memory from a plurality of the processors arranged in the same column is input from the external memory to a plurality of the processors arranged in the same row.
[0124] (Aspect 8) The arithmetic device according to any one of Aspects 1 to 6, wherein data output from the plurality of processors to the external memory is input from the external memory to the processors such that one processor holds a plurality of pieces of data.
[0125] (Aspect 9) The arithmetic device according to any one of Aspects 1 to 6, wherein data output from the plurality of processors to the external memory is copied and input from the external memory to the processors so that the plurality of processors hold the same data.
[0126] (Aspect 10) A computing device according to any one of Aspects 1 to 6, wherein data formed into a multidimensional array by a plurality of groups (21) formed of two-dimensionally arranged data is stored in an external memory provided outside the computing device, and after inputting data in the nth row and mth column of the plurality of groups from the external memory to the processor in the nth row and mth column, the processor outputs the data to the corresponding external memory, and the external memory inputs data to the plurality of processors so that the arrangement of the data formed into a multidimensional array by the plurality of groups becomes a two-dimensional array.
[0127] (Aspect 11) An arithmetic device according to any one of Aspects 1 to 6, wherein one or more of the external memories are associated with the processors on a row or column basis, and the processor located at one end of the row or column outputs data to the external memory, while the other processors move data toward the processor at the one end, and data is input from the external memory to the processor at the other end.
[0128] (Aspect 12) The arithmetic device according to any one of aspects 1 to 11, wherein data is transferred between a plurality of the external memories.
[0129] (Aspect 13) A data movement method for a computing device having a plurality of external memories and a plurality of processors, wherein the processors are associated with the external memories that input and output data, the external memories store data in association with the processors that input and output data, and data is input and output between the processors and the external memories.
Claims
1. A computing device (10) comprising a plurality of external memories (14) and a plurality of processors (12), wherein the processors are associated with the external memories that input and output data, and the external memories store data in association with the processors that input and output data.
2. The computing device according to claim 1, wherein data transferred from an external memory (20) provided outside the computing device to the external memory is sorted in accordance with the location of the processor and stored in the external memory.
3. The computing device according to claim 1 or 2, wherein the external memory is arranged in a plurality of different directions relative to the processor.
4. The computing device according to claim 1 or 2, wherein the external memory is arranged in a one-to-one correspondence with the processor.
5. The arithmetic device according to claim 1 or 2, wherein the plurality of processors are divided into a plurality of groups (30) in one row or one column, and the external memory is arranged for each of the groups.
6. The arithmetic device according to claim 1 or 2, wherein the plurality of processors are divided into a plurality of groups in at least one of a plurality of rows and a plurality of columns, and the external memory is arranged for each of the groups.
7. The arithmetic device according to claim 1 or claim 2, wherein data output to the external memory from a plurality of the processors arranged in the same row is input from the external memory to a plurality of the processors arranged in the same column, and data output to the external memory from a plurality of the processors arranged in the same column is input from the external memory to a plurality of the processors arranged in the same row.
8. The arithmetic device according to claim 1 or 2, wherein data output from a plurality of said processors to said external memory is input from said external memory to said processors so that one said processor holds a plurality of data.
9. The arithmetic device according to claim 1 or 2, wherein data output from the plurality of processors to the external memory is copied and input from the external memory to the processors so that the plurality of processors hold the same data.
10. The arithmetic device according to claim 1 or 2, wherein data arranged in a multidimensional array by a plurality of groups (21) formed of two-dimensionally arranged data is stored in an external memory provided outside the arithmetic device, and after inputting data in the nth row and mth column of the plurality of groups from the external memory to the processor in the nth row and mth column, the processor outputs the data to the corresponding external memory, and the external memory inputs data to the plurality of processors so that the arrangement of the data arranged in the multidimensional array by the plurality of groups becomes a two-dimensional array.
11. The arithmetic device according to claim 1 or claim 2, wherein one or more of the external memories are associated with the processors in each row or column, and the processor located at one end of the row or column outputs data to the external memory, while the other processors move data toward the processor at the one end, and data is input from the external memory to the processor at the other end.
12. The arithmetic device according to claim 1 or 2, wherein data is transferred between a plurality of said external memories.
13. A data movement method for a computing device having a plurality of external memories and a plurality of processors, wherein the processors are associated with the external memories that input and output data, the external memories store data in association with the processors that input and output data, and data is input and output between the processors and the external memories.
Citation Information
Patent Citations
Pruning method and device of neural network, equipment and medium
CN114662689A
Parallel computer and compiler
JP1994028324A
Batch processing method for data transfer
JP1995093269A
Floating point multiply-accumulate unit for deep learning
US20220188075A1