Computation device and data transfer method
The computing device addresses the need for reduced memory capacity by enabling simultaneous data movement between processors, optimizing memory usage and processing efficiency without external memory.
Patent Information
- Application Number
- PCT/JP2025/023110
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2025-06-26
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional arithmetic devices require large external memories to accommodate increasing data processing needs, which can be inefficient and costly.
A computing device with multiple processors that share internal memory, allowing data to be moved simultaneously between processors without the need for external memory, utilizing a pipeline and symmetric diffusion to reduce memory capacity requirements.
The solution enables efficient data movement within the device, reducing the need for external memory and enhancing processing capabilities by minimizing internal memory usage while maintaining data processing efficiency.
Smart Images

Figure JP2025023110_12022026_PF_FP_ABST
Abstract
Description
Arithmetic device and data movement method CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application is based on Japanese Application No. 2024-130514, filed on August 7, 2024, the contents of which are incorporated herein by reference.
[0002] The present disclosure relates to a computing device and a data movement method.
[0003] 2. Description of the Related Art Arithmetic units have been developed that are configured with a plurality of processors arranged two-dimensionally (for example, in an array or a mesh structure).
[0004] As such a computing device, for example, Patent Document 1 describes a computing device including a first plurality of processing cores arranged in an array and a second plurality of processing cores arranged in an array.
[0005] U.S. Pat. No. 1,086,574
[0006] 11 , a conventional arithmetic device 100 includes a plurality of processors (Processing Elements, hereinafter referred to as "PEs") 102 and an external memory 104. Arrows in FIG. 11 indicate examples of data input / output between the PEs 102 and between the PEs 102 and the external memory 104.
[0007] In recent years, the capacity of the external memory 104 has tended to increase in order to accommodate various data to be processed by the arithmetic device 100. In order to reduce the capacity of the external memory 104, it is conceivable to store multiple pieces of data to be processed in advance in the internal memory of each PE 102. However, this approach requires increasing the capacity of the internal memory of the PE 102.
[0008] An object of the present invention is to provide a computing device and a data movement method that do not require an external memory or that can reduce the capacity of the external memory.
[0009] One embodiment of the present invention is a computing device comprising a plurality of processors each having an internal memory, wherein the processor outputs retained data, which is data retained in the internal memory, to a plurality of other processors simultaneously as movement data, and the internal memory retains the movement data input from the other processors in a manner that makes it possible to distinguish between the retained data and the movement data.
[0010] According to this configuration, data used for calculations in a processor is input as transfer data from another processor, so the calculation device does not require an external memory, or the capacity of the external memory can be reduced.
[0011] In the above-described arithmetic device, the data to be moved may be output from the processor to another processor via a pipeline. With this configuration, multiple pieces of data to be moved are simultaneously moved between different processors within the arithmetic device, thereby shortening the time it takes for multiple pieces of data to move to a desired processor.
[0012] In the above-mentioned computing device, the multiple processors may be arranged in a one-dimensional or two-dimensional array, and the movement data may be output from the processors to a central processor located at the center of the one-dimensional or two-dimensional array, which then outputs the input movement data to the multiple other processors. With this configuration, the movement data output to the central processor is diffused from the central processor to the other processors. This allows the movement data to be supplied to the other processors more quickly, regardless of the location of the processor that outputs the movement data.
[0013] In the above-mentioned computing device, the output of the moving data from the central processor to the other processors may be performed symmetrically around the central processor. According to this configuration, the processing for diffusing the moving data from the central processor to the other processors can be simplified.
[0014] In the above-mentioned computing device, the plurality of processors may be grouped into a plurality of processor groups, and each of the processor groups may have a representative processor to which the movement data is input from the central processor, and the representative processor may output the movement data to the other processors constituting the processor group to which it belongs. With this configuration, the movement data input from the central processor to the plurality of representative processors is diffused from the representative processor to the other processors constituting the processor group to which it belongs, thereby reducing the movement cycle of the movement data.
[0015] In the above-described computing device, the plurality of processor groups may be arranged symmetrically around the central processor. With this configuration, it is possible to simplify the process of diffusing the movement data from the central processor to the other processors.
[0016] In the above-described computing device, when the transfer data needs to be transferred multiple times from the processor to the central processor, the processor that temporarily stores the transfer data for each transfer may be set as an intermediate processor. With this configuration, the transfer data is efficiently output to the central processor.
[0017] In the above-described computing device, the processor may hold at most one of the pieces of movement data being moved to another of the processors. With this configuration, the processor can minimize the internal memory used for supplying the movement data.
[0018] In the above-mentioned computing device, when new movement data is input, the representative processor may transfer the movement data that it has held to another adjacent processor in the same processor group, and when new movement data is input, the other processor may transfer the movement data that it has held to another adjacent processor in the same processor group. With this configuration, the processor can minimize the amount of internal memory used in supplying movement data.
[0019] In the above-mentioned computing device, the processor that outputs the movement data to the central processor may be preset as an output processor, and the other processors may move their own retained data toward the output processor along a predetermined path in accordance with the output of the movement data by the output processor to the central processor. With this configuration, the movement data always moves from the output processor to the central processor, so the movement path of the movement data is the same, and the movement process of the movement data can be simplified.
[0020] A data movement method of one aspect of the present invention is a data movement method for a computing device having multiple processors each having an internal memory, in which the processor outputs retained data, which is data stored in the internal memory, to multiple other processors simultaneously as movement data, and the internal memory retains the movement data input from the other processors in a manner that makes it possible to distinguish between the retained data and the movement data.
[0021] According to the present invention, external memory is not required, or the capacity of the external memory can be reduced.
[0022] The above and other objects, features, and advantages of the present disclosure will become more apparent from the following detailed description taken in conjunction with the accompanying drawings. The drawings include: FIG. 1 is a schematic diagram of an accelerator according to an embodiment; FIG. 2 is a schematic diagram illustrating data transfer from a processing element (PE) to another processing element (PE) according to an embodiment, where (A) shows data held in advance in the internal memory of each processing element (PE); (B) shows data after the first data transfer; (C) shows data after the second data transfer; and (D) shows data after the sixth data transfer. FIG. 3 is a schematic diagram illustrating the arrangement of processing elements (PEs) constituting the accelerator according to an embodiment; FIG. 4 is a schematic diagram illustrating data transfer after the first cycle in the accelerator of FIG. 3; FIG. 5 is a schematic diagram illustrating data transfer after the second cycle in the accelerator of FIG. 3; FIG. 6 is a schematic diagram illustrating data transfer after the third cycle in the accelerator of FIG. 3; and FIG. 7 is a schematic diagram illustrating data transfer after the fourth cycle in the accelerator of FIG. 3. Fig. 8 is a schematic diagram of a form in which data is moved while shifting data held in PEs according to an embodiment. Fig. 9 is a schematic diagram showing the arrangement of PEs in a form in which data is moved while shifting data held in PEs according to an embodiment. Fig. 10 is a schematic diagram of a form in which data is moved between multiple systems. Fig. 11 is a schematic diagram of a conventional arithmetic device.
[0023] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. Note that the embodiments described below are examples of how the present disclosure may be implemented, and the present disclosure is not limited to the specific configurations described below. When implementing the present disclosure, specific configurations according to the embodiments may be appropriately adopted.
[0024] FIG. 1 is a schematic diagram of an accelerator 10, which is a computing device according to this embodiment.
[0025] The accelerator 10 is composed of a plurality of PEs 12. In the example of Fig. 1, a total of 16 PEs 12 are arranged two-dimensionally in four rows and four columns, but the accelerator 10 may include any number of PEs 12. In the following description, the coordinates of the upper left PE 12 are set to (0, 0), and each PE 12 is also distinguished by its coordinates.
[0026] The PEs 12 constituting the accelerator 10 are electrically connected by wiring 14 for data movement between the PEs 12. A PE 12 can move data between at least other PEs 12 adjacent to it on the top, bottom, left, and right sides via the wiring 14. In this embodiment, data movement is a concept that also includes copying data, and can also be expressed as input, output, and supply of data. Note that although the wiring 14 is omitted in other figures, even in figures where the wiring 14 is omitted, adjacent PEs 12 are still connected by the wiring 14. Furthermore, two or more wirings 14 may be used to connect the PEs 12 to enable simultaneous input and output of multiple data between two PEs 12.
[0027] The PE 12 of this embodiment includes an arithmetic circuit 16 and an internal memory 18 .
[0028] The arithmetic circuit 16 uses data stored in the internal memory 18 to perform various arithmetic operations such as ==, !=, >, >=, <, <=, >>, <<, or, and, min, max, clip, add, sub, mul, div, mod, macc, etc.
[0029] The internal memory 18 is a storage unit that holds data. The arithmetic circuit 16 performs arithmetic processing using the data held in the internal memory 18, and the arithmetic results are held in the internal memory 18 or output to a memory outside the accelerator 10. Note that before the accelerator 10 starts an arithmetic operation, data is input to the internal memory 18 from a memory outside the accelerator 10. Data that is input to each PE 12 from a memory outside the accelerator 10 and held in the internal memory 18 is called held data.
[0030] Here, the PE 12 of this embodiment outputs the retained data stored in the internal memory 18 to multiple other PEs 12 simultaneously as movement data. Note that in this embodiment, as an example, the PE 12 outputs the retained data to the other PEs 12 as movement data. The internal memory 18 then retains the movement data input from the other PEs 12 so that the retained data can be distinguished from the movement data. Note that the movement data is, as an example, a copy of the retained data. The same data is passed to each PE 12 as movement data. Each PE 12 performs calculations using the passed movement data, and discards or renders the data unnecessary after the calculation. Therefore, each PE 12 does not retain the movement data as retention data, so even if movement data is passed, the impact on the internal memory 18 is minimal.
[0031] 2A and 2B are schematic diagrams showing data transfer from a PE 12 to another PE 12 in this embodiment. In Fig. 2A, the numbers in the PE 12 indicate data held in the internal memory 18 of each PE 12. Similarly, in the following figures, the numbers in the PE 12 indicate data held in the internal memory 18 of each PE 12.
[0032] In the example of Figure 2 (A), the internal memory 18 of each PE 12 holds different held data, such as the held data of the PE 12 at coordinates (0,0) being "00", the held data of the PE 12 at coordinates (0,1) being "01", and the held data of the PE 12 at coordinates (0,3) being "03".
[0033] 2B shows the state after the first data transfer, in which the data "00" held in the PE 12 at coordinates (0,0) is transferred to all other PEs 12. The other PEs 12 store the data "00" transferred from coordinates (0,0) in their internal memories 18, distinguishing it from their own data.
[0034] 2C shows the state after the second data movement, in which "01", the data held by the PE 12 at coordinates (0, 1), is moved to all other PEs 12. Furthermore, FIG. 2D shows the state after the sixth data movement, in which "11", the data held by the PE 12 at coordinates (1, 1), is moved to all other PEs 12. As an example, when new movement data is input to the PE 12, the PE 12 overwrites the movement data that it has been holding up until then and stores the new movement data in the internal memory 18.
[0035] For example, each time data to be moved is transferred, each PE 12 performs an operation using the data to be moved and the data held by itself in the operation circuit 16. The operation result by the operation circuit 16 is held in its own internal memory 18 and used for the next operation, or is output to a memory outside the accelerator 10, another accelerator 10, etc.
[0036] Note that the PE 12 does not need to move the movement data to all of the other PEs 12 that make up the accelerator 10. For example, the movement data may be moved to some of the other PEs 12 that make up the accelerator 10. Also, different movement data may be used for each row, and the movement data may be moved to the same row. Similarly, different movement data may be used for each column, and the movement data may be moved to the same column.
[0037] In this way, in the accelerator 10 of this embodiment, data used for calculations in a PE 12 is transferred from other PEs 12, so the accelerator 10 does not require an external memory as conventionally provided, or even if the accelerator 10 is provided with a conventional external memory, the capacity can be reduced. Furthermore, since the internal memory 18 provided in the PE 12 is used to hold data held by other PEs 12, by increasing the capacity of the internal memory 18 instead of the external memory, the accelerator 10 can hold more data and make processing easier.
[0038] (Diffusion and Movement Processing of Movement Data) Next, the diffusion and movement of movement data that is performed when the accelerator 10 includes a larger number of PEs 12 will be described with reference to FIGS.
[0039] 3 is a schematic diagram showing the arrangement of PEs 12 constituting the accelerator 10 of this embodiment. In the example of FIG. 3, the accelerator 10 is a group of PEs composed of 16 PEs in the row direction and 16 PEs in the column direction, for a total of 256 PEs 12. Note that in FIG. 3 and other figures, some of the PEs 12 constituting the accelerator 10 are omitted.
[0040] In the example of Figure 3, the data held by PE12 at coordinates (0,0) is "00", the data held by PE12 at coordinates (0,5) is "05", and the data held by PE12 at coordinates (0,9) is "09". As an example, each PE12 holds different held data.
[0041] When the accelerator 10 includes a large number of PEs 12, a single movement, in other words, moving data one by one in one cycle of the arithmetic processing of the PEs 12, takes time and is inefficient. Therefore, the accelerator 10 of this embodiment uses a pipeline to output the movement data from one PE 12 to another PE 12. Parallel processing using the pipeline allows multiple movement data to move simultaneously between different PEs 12, thereby shortening the time it takes for multiple movement data to move to the desired PE 12.
[0042] Furthermore, the accelerator 10 of this embodiment outputs movement data from the PEs 12 toward a central PE 12C, which is the PE 12 located at the center of the two-dimensional array. The central PE 12C then outputs the input movement data to multiple other PEs 12. In the following description, this movement of movement data is referred to as a diffusion movement process.
[0043] The diffusion and movement process diffuses movement data output from a PE 12 to a central PE 12C from the central PE 12C to other PEs 12, thereby enabling the movement data to be supplied quickly to other PEs 12 regardless of the location of the PE 12 that outputs the movement data. Furthermore, the diffusion and movement process differs in the direction in which the movement data moves toward the central PE 12C from the direction in which the movement data is diffused from the central PE 12C toward other PEs 12, thereby preventing the movement paths of the movement data from overlapping within one cycle and reducing the number of other PEs 12 through which the movement data passes.
[0044] Furthermore, in the diffusion migration process, the multiple PEs 12 that make up the accelerator 10 are grouped into multiple sub-PE groups 20. A representative PE 12A is set in each sub-PE group 20, to which migration data is input from a central PE 12C. The representative PE 12A outputs the migration data to the other PEs 12 that make up the sub-PE group 20 that includes the representative PE 12A. In this way, the diffusion migration process of this embodiment outputs the migration data that has been moved to the central PE 12C to the multiple representative PEs 12A, and then diffuses the migration data from the representative PE 12A to the other PEs 12 that make up the sub-PE group 20 that includes the representative PE 12A, thereby reducing the migration cycle of the migration data.
[0045] Furthermore, the output of the movement data from the central PE 12C to the other PEs 12 is performed symmetrically with respect to the central PE 12C. This simplifies the process of diffusing the movement data from the central PE 12C to the other PEs 12. Note that symmetry refers to point symmetry or line symmetry.
[0046] Therefore, the plurality of sub PE groups 20 are arranged symmetrically with respect to the central PE 12 C. With this configuration, the process for diffusing the movement data from the central PE 12 C to the other PEs 12 can be simplified.
[0047] Next, the details of the diffusion and migration process will be described with reference to FIGS. 3 to 7. In the examples of FIGS. 3 to 7, as described above, the accelerator 10 is configured with 16 PEs 12 in the row direction and 16 PEs in the column direction, for a total of 256 PEs 12. The distance that the migration data can move in one cycle of the PE 12's arithmetic processing is assumed to be a maximum of eight PEs 12. In this example, one cycle of migration is described as one cycle of the PE 12's arithmetic processing, but the cycles for migration and arithmetic processing do not necessarily have to be the same. For example, one cycle of migration may be as long as two arithmetic processing cycles, or conversely, may be as short as 0.5 arithmetic processing cycles. However, because the migration data is data that is used in a calculation immediately after the diffusion and migration and is not retained but discarded, it is desirable to set the cycle so that the next migration data is obtained at the timing when the next calculation is possible.
[0048] The number of sub PE groups 20 is determined according to the maximum movement distance of the movement data in one cycle, and in this embodiment, since the maximum movement distance is 8, it is (8 / 2) x (8 / 2), that is, a group of 4 rows and 4 columns of PEs 12. Note that if the maximum movement distance is an odd number, it cannot be divided by 2, so the number of groups is determined as the value obtained by subtracting 1 from the maximum movement distance and multiplying this number by 1 / 2.
[0049] 4 to 7, numbers enclosed in circles indicate data being moved from one PE 12 to another PE 12. The PE 12 indicated by horizontal hatching is the center PE 12C. As described above, the sub PE group 20 is arranged symmetrically with the center PE 12C as the center. Therefore, in this embodiment, the center PE 12C is the four PEs 12 located at the center of the 16 columns and 16 rows of PEs 12. In other words, the PEs 12 with coordinates (7,7), (7,8), (8,7), and (8,8) are set as the center PE 12C.
[0050] The PEs 12 indicated by vertical hatching are representative PEs 12A. One representative PE 12A is set for each sub-PE group 20. The position of the PE 12 set as the representative PE 12A is, for example, the position closest to the center PE 12C within the sub-PE group 20. Note that in the sub-PE group 20 in which the center PE 12C is set, a PE 12 different from the center PE 12C is set as the representative PE 12A.
[0051] The PEs 12 indicated by diagonal hatching are transit PEs 12B. The transit PEs 12B are set as PEs 12 that temporarily store the movement data when the movement data moves from the PE 12 that is the output source of the movement data. That is, in the diffusion movement process, when multiple cycles are required to output the movement data from the PE 12 to the center PE 12C, a transit PE 12B is set to temporarily store the movement data for each cycle. This allows the movement data to be efficiently output to the center PE 12C. Note that the position of the transit PE 12B shown in the figure is an example, and the transit PE 12B may be set at another position.
[0052] The thin solid arrows a indicate the path of data movement from each PE 12 to the central PE 12C. The thick solid arrows b indicate the path of data movement when data movement is output from the central PE 12C to other PEs 12, in other words, when data is diffused. That is, the arrows b indicate the path of data movement from the central PE 12C to the representative PE 12A.
[0053] The dashed bold arrow c indicates the migration path when diffusing migration data from the representative PE 12A in the accelerator 10 in which a sub-PE group 20 exists further outside the 4-row, 4-column sub-PE group 20.
[0054] The thin dashed arrow d points in the opposite direction to the arrow a and indicates a movement path when movement data is temporarily saved within the same sub-PE group 20. The processing of movement data in the direction of the arrow d will be described in detail later.
[0055] 4 shows the data after the first cycle of data movement, in which the PE 12 at coordinates (0, 0) copies the stored data "00" as movement data. The movement data "00" moves to the transit PE 12B at coordinates (3, 5).
[0056] 5 shows the state after the second cycle of data movement. The moved data "00" moves from the transit PE 12B to the center PE 12C at coordinates (7,7). Note that since the maximum movement distance is 8, the moved data "00" can also move to coordinates (7,8), (8,7), and (8,8) at the same time, and the moved data "00" is copied to the three other center PEs 12C.
[0057] In the second cycle, the data "01" held in the PE 12 at the coordinates (0, 1) is copied by the pipeline as the moving data, and moves to the transit PE 12B at the coordinates (3, 5).
[0058] 6 shows the state after the data movement in the third cycle. The movement data "00" is copied and diffused from the central PE 12C to the representative PE 12A of the other sub-PE group 20. As described above, the distance that data can move in one cycle is a maximum of eight PEs 12, so the movement data moves from the central PE 12C to the representative PE 12A shown in FIG. 6 in one cycle.
[0059] Furthermore, the moving data "01" is moved and copied from the transit PE 12B to the central PE 12C by the pipeline. Note that, in the sub PE group 20 in which the central PE 12C is set, the moving data "01" is moved to the central PE 12C, so the central PE 12C located closest to the central PE 12C cannot be set as the representative PE 12A. Therefore, the PE 12 one position outside the central position is set as the representative PE 12A, and the moving data "00" held by the central PE 12C is moved to the representative PE 12A of the sub PE group 20. Note that in the following description, the moving of the moving data previously held to another PE 20 in conjunction with the movement of new moving data to a PE 20 is also referred to as "saving."
[0060] As in the data movement in the third cycle, the direction in which the moving data "00" is diffused is opposite to the direction in which the moving data "01" moves to the central PE 12C. For this reason, the accelerator 10 has a circuit configuration that allows moving data to be moved in opposite directions at the same time. This circuit configuration is, for example, a configuration in which two wires 14 are provided between two PEs 12.
[0061] Furthermore, in the third cycle, the moving data "02" moves from the PE 12 at the coordinates (0, 2) to the transit PE 12B at the coordinates (3, 5) through the pipeline.
[0062] 7 shows the data after the fourth cycle of data movement. The moved data “00” moved to each representative PE 12A is copied and diffused by simultaneously outputting it to all other PEs 12 constituting the sub-PE group 20 including the representative PE 12A, and is used in the calculations by each PE 20.
[0063] Furthermore, the moving data "01" moves from the central PE 12C to the representative PE 12A of another adjacent sub-PE group 20. The moving data "01" moved to the representative PE 12A is simultaneously copied and output to all other PEs 12 constituting each sub-PE group 20 in the next cycle, the fifth cycle.
[0064] Furthermore, the moving data "02" is moved and copied from the transit PE 12B to the central PE 12C. Note that, in the sub PE group 20 in which the central PE 12C is set, since the moving data "02" moves to the central PE 12C, the PE 12 immediately outside the central position is set as the representative PE 12A, and the moving data "01" held by the central PE 12C is moved to the representative PE 12A of the sub PE group 20. Note that the moving data "00" held by the representative PE 12A is saved to another PE 12 adjacent to the representative PE 12A along arrow d in order to be held in that PE 12.
[0065] In this way, the PE 12 of this embodiment holds a maximum of one piece of movement data being moved to another PE 12. Therefore, when new movement data is input, the PE 12 moves the movement data that it has held up until then to another adjacent PE 12. That is, in the fourth cycle, movement data "01" moves to the representative PE 12A, and the movement data "00" held in the representative PE 12A is saved to another adjacent PE 12 included in the same sub-PE group 20. When new movement data is further moved to the sub-PE group 20, the movement data is saved to another adjacent PE 12 included in the same sub-PE group 20 in turn along arrow d. Arrow d indicates the save path of such movement data.
[0066] By saving the movement data indicated by the arrow d, the PE 12 can minimize the capacity of the internal memory 18 used to supply movement data. Furthermore, each movement data moved to each sub-PE group 20 is held by one of the PEs 12 constituting the sub-PE group 20. Once the same movement data is held in all of the sub-PE groups 20, the movement data is used to prepare for arithmetic processing using the same data. Then, when movement data moved in the previous cycle becomes necessary for arithmetic processing, the PE 12 holding the movement data outputs the movement data to all other PEs 12 constituting the sub-PE group 20 to which it belongs, so that the movement data can be used for arithmetic processing. Therefore, as can be seen from FIG. 7 , as the cycle of moving movement data progresses, the movement data saved increases for the sub-PE groups 20 closer to the center PE 12C.
[0067] Furthermore, in the fourth cycle, if there is a sub-PE group 20 further outside than the illustrated sub-PE group 20, the movement data "00" is output to the representative PE 12A of the outer sub-PE group 20, as indicated by arrow c. Furthermore, in the fourth cycle, the movement data "03" is moved from the PE 12 at coordinates (0, 3) to the transit PE 12B at coordinates (3, 5). The diffusion movement processing of this embodiment diffuses the movement data output from the other PEs 12 in the same manner as the diffusion of movement data described with reference to FIGS. 3 to 7 .
[0068] In addition, when a large amount of movement data is used in a calculation or when the movement data is to be immediately reused, it is desirable to have a large amount of movement data in movement. In this case, one movement data in movement in each PE 12 may not be enough. In this case, the amount of movement data available in each PE 12 may be increased from one to two or more.
[0069] (Diffusion and Movement Processing While Shifting Retained Data) Next, with reference to Figures 8 and 9, a form of moving data to the central PE 12C while shifting the retained data held by the PE 12 will be described. Figure 8 shows the state before the data movement, and is the same as the arrangement of the PE 12 and retained data shown in Figure 3. Figure 9 shows the arrangement of data after the data movement in the first cycle.
[0070] In the accelerator 10 of this embodiment, the order in which the migration data is moved is determined in advance, and the retained data is shifted from one PE 12 to another PE 12 in accordance with this order. That is, in the above-described diffusion migration process, the retained data of the PE 12 is copied and used as the migration data, and therefore the retained data is not moved, whereas in this embodiment, the retained data of the PE 12 itself is moved to another PE 12. However, if there is a possibility that the retained data will be reused, the retained data may be saved or a copy of the retained data may be used in the diffusion migration process.
[0071] For this reason, the PE 12 that outputs movement data to the center PE 12C is set in advance as the output PE 12D. In the examples of Figures 8 and 9, the PE 12 at coordinates (0, 0) is set as the output PE 12D. Then, in accordance with the output of movement data to the center PE 12C by the output PE 12D, the other PEs 12 move their own held data toward the output PE 12D along a predetermined path.
[0072] 9, the held data at coordinates (0,0) of the output PE 12D is moved as moving data to the PE 12 at coordinates (3,5) of the transit PE 12B. Then, the held data "01" of the PE 12 at adjacent coordinates (0,1) is moved to the output PE 12D. Similarly, the held data of each PE 12 moves in turn toward the output PE 12D in a predetermined order.
[0073] In the example of FIG. 9 , as indicated by the arrow e superimposed on the PEs 12, the retained data of the PEs 12 in odd-numbered rows, such as the first row, moves from right to left. To the rightmost PE 12 at the end of the odd-numbered rows, the retained data moves from the rightmost PE 12 immediately adjacent below, i.e., in the row below. Furthermore, the retained data of the PEs 12 in even-numbered rows, such as the second row, moves from left to right. To the leftmost PE 12 at the beginning of the even-numbered rows, the retained data moves from the leftmost PE 12 immediately adjacent below, i.e., in the row below. Note that the retained data movement path indicated by the arrow e is merely an example and is not limited to this. For example, data may be moved from right to left only in the first row, and after all the data in the first row has been moved, the second and subsequent rows may be shifted up by one row, and then only data in the first row may be moved from right to left again, and this movement may be repeated.
[0074] In the second cycle of movement, the movement data "00" is moved from the transit PE 12B to the center PE 12C. The held data "01" of the output PE 12D is moved to the transit PE 12B, and the held data "02" of the PE 12 adjacent to the right is moved to the output PE 12D. Then, the held data of the other PEs 12 are moved in turn toward the output PE 12D as described above. Note that the movement and diffusion of the movement data thereafter is the same as the diffusion movement processing described with reference to FIGS. 4 to 7.
[0075] This process simplifies the process of moving data because the path of movement of the data from the output PE 12D to the center PE 12C is always the same. The center PE 12C may initially arrange the retained data so that the data to be diffused is the data to be moved, and then move this retained data. Conversely, the initial position of the data to be diffused may be the center PE 12C. This reduces the processing time required to move the initial data to be diffused to the center PE 12C.
[0076] Although the present disclosure has been described using the above-mentioned embodiments, the technical scope of the present disclosure is not limited to the scope described in the above-mentioned embodiments. Various modifications or improvements can be made to the above-mentioned embodiments without departing from the gist of the disclosure, and such modifications or improvements are also included in the technical scope of the present disclosure.
[0077] In the above embodiment, the diffusion and movement processing is performed on a plurality of PEs 12 arranged in a two-dimensional array. However, the present disclosure is not limited to this. For example, the diffusion and movement processing may be performed on a plurality of PEs 12 arranged in a one-dimensional array. In this case, when different movement data are used for each row and the movement data are moved to the same row, the central PE 12C is located at the center of the row and the movement data are aligned vertically. Similarly, when different movement data are used for each column and the movement data are moved to the same column, the central PE 12C is located at the center of the column and the movement data are aligned horizontally. This case is similar to the case where the diffusion and movement processing is performed on a plurality of PEs 12 arranged in a two-dimensional array. That is, the process of diffusing the same movement data to a one-dimensional array in the same row and diffusing another same movement data to a one-dimensional array in another row is similarly performed, and the process of diffusing the same movement data to a one-dimensional array in the same column and diffusing another same movement data to a one-dimensional array in another row is similarly performed.
[0078] In the above embodiment, a configuration in which a series of diffusion and migration processes are performed across the entire accelerator 10 has been described. However, the present disclosure is not limited to this. For example, the multiple PEs 12 constituting the accelerator 10 may be divided into multiple processing regions, and data required for each processing region may be allocated to multiple PEs 12, and the diffusion and migration process may be performed for each processing region. The multiple PEs 12 included in a processing region may be divided into multiple sub-PE groups 20. This allows different migration data to be handled for each processing region, enabling calculations to be performed using different data for each processing region. For example, a row of PEs 12 may be divided into a single processing region, and each row may be divided into a separate processing region, and the diffusion and migration process may be performed for each row of the processing region. Alternatively, a column of PEs 12 may be divided into a single processing region, and each column may be divided into a separate processing region, and the diffusion and migration process may be performed for each column of the processing region. Alternatively, n rows and m columns of PEs 12 may be divided into a single processing region, and the diffusion and migration process may be performed for each column of the processing region.
[0079] In the above embodiment, the diffusion and migration process is performed within one accelerator 10, i.e., one system. However, the present disclosure is not limited to this. For example, as shown in FIG. 10 , the diffusion and migration process may be performed in a configuration using multiple accelerators 10 as one system. In this configuration, data moves toward the hatched area in FIG. 10 , which is the center of the multiple accelerators 10, and then diffuses from there. This eliminates the need to store data for each accelerator 10, and allows the same data to be diffused to all PEs 12.
[0080] In the above embodiment, the sub PE group 20 is described as being square, but the present disclosure is not limited to this. The sub PE group 20 may be, for example, one row or one column, a rectangle, or another rectangular shape.
[0081] In the above embodiment, a configuration has been described in which the central PE 12C is set at the center of the plurality of PEs 12 that constitute the accelerator 10. However, the present disclosure is not limited to this. The central PE 12C may be set to a PE 12 other than the center of the plurality of PEs 12 that constitute the accelerator 10, and the movement data may be diffused from that PE 12.
[0082] In the above embodiment, the accelerator 10 does not include an external memory, but the present disclosure is not limited to this. The accelerator 10 may include an external memory having a minimum required memory area.
[0083] Next, the features of the present disclosure are as follows.
[0084] (Aspect 1) A computing device (10) comprising a plurality of processors (12) each having an internal memory (18), wherein the processor outputs retained data, which is data retained in the internal memory, to a plurality of other processors simultaneously as movement data, and the internal memory retains the movement data input from the other processors in a manner that allows the retained data to be distinguished from the movement data.
[0085] (Aspect 2) The arithmetic device according to aspect 1, wherein the movement data is output from the processor to the other processor via a pipeline.
[0086] (Aspect 3) The computing device according to aspect 1 or aspect 2, wherein the plurality of processors are arranged in a one-dimensional or two-dimensional array, the movement data is output from the processors to a central processor (12C) that is the processor located at the center of the one-dimensional or two-dimensional array, and the central processor outputs the input movement data to the plurality of other processors.
[0087] (Aspect 4) The arithmetic device according to aspect 3, wherein the output of the movement data from the central processor to the other processors is performed symmetrically with respect to the central processor.
[0088] (Aspect 5) The computing device according to aspect 3 or aspect 4, wherein the plurality of processors are grouped into a plurality of processor groups (20), the processor groups are each set up with a representative processor (12A) to which the movement data is input from the central processor, and the representative processor outputs the movement data to the other processors that make up the processor group to which it belongs.
[0089] (Aspect 6) The arithmetic device according to aspect 5, wherein the plurality of processor groups are arranged symmetrically with the central processor at the center.
[0090] (Aspect 7) The arithmetic device according to any one of aspects 3 to 6, wherein, when multiple transfers are required to output the transfer data from the processor to the central processor, the processor that temporarily stores the transfer data for each transfer is set as a relay processor (12B).
[0091] (Aspect 8) The arithmetic device according to any one of aspects 3 to 7, wherein the processor holds a maximum of one piece of the movement data that is being moved to another of the processors.
[0092] (Aspect 9) The computing device according to aspect 5, wherein, when new movement data is input, the representative processor moves the movement data that it had been holding to another adjacent processor that is included in the same processor group, and when new movement data is input, the other processor moves the movement data that it had been holding to another adjacent processor that is included in the same processor group.
[0093] (Aspect 10) A computing device according to any one of aspects 3 to 9, wherein the processor that outputs the movement data to the central processor is preset as an output processor (20D), and the other processors move their own retained data towards the output processor along a predetermined route in accordance with the output of the movement data by the output processor to the central processor.
[0094] (Aspect 11) A data transfer method for a computing device having a plurality of processors each having an internal memory, wherein the processor outputs retained data, which is data retained in the internal memory, to a plurality of other processors simultaneously as transfer data, and the internal memory retains the transfer data input from the other processors in a manner that allows the retained data to be distinguished from the transfer data.
Claims
1. A computing device (10) comprising a plurality of processors (12) each having an internal memory (18), wherein the processor outputs stored data, which is data stored in the internal memory, to a plurality of other processors simultaneously as movement data, and the internal memory stores the movement data input from the other processors in a manner that makes it possible to distinguish between the stored data and the movement data.
2. The arithmetic device according to claim 1, wherein the output of the movement data from the processor to the other processor is performed by a pipeline.
3. The arithmetic device according to claim 1 or claim 2, wherein the plurality of processors are arranged in a one-dimensional or two-dimensional array, the movement data is output from the processors to a central processor (12C) that is the processor located at the center of the one-dimensional or two-dimensional array, and the central processor outputs the input movement data to the plurality of other processors.
4. The arithmetic device according to claim 3, wherein the output of the movement data from the central processor to the other processors is performed symmetrically with respect to the central processor.
5. The computing device according to claim 3, wherein the plurality of processors are grouped into a plurality of processor groups (20), and the processor groups are each assigned a representative processor (12A) to which the movement data is input from the central processor, and the representative processor outputs the movement data to the other processors that make up the processor group to which it belongs.
6. The computing device according to claim 5, wherein the plurality of processor groups are arranged symmetrically with the central processor at the center.
7. The computing device according to claim 3, wherein when the output of the movement data from the processor to the central processor requires multiple movements, the processor that temporarily holds the movement data for each movement is set as a transit processor (12B).
8. The computing device according to claim 3, wherein the processor holds at most one of the moving data items in the process of moving to another of the processors.
9. The arithmetic device of claim 5, wherein, when new movement data is input, the representative processor moves the movement data it has been holding to another adjacent processor in the same processor group, and when new movement data is input, the other processor moves the movement data it has been holding to another adjacent processor in the same processor group.
10. The computing device according to claim 3, wherein the processor that outputs the movement data to the central processor is preset as an output processor (20D), and the other processors move their own retained data toward the output processor along a predetermined path in accordance with the output of the movement data by the output processor to the central processor.
11. A data transfer method for an arithmetic device having multiple processors each having an internal memory, wherein the processor outputs retained data, which is data retained in the internal memory, to multiple other processors simultaneously as transfer data, and the internal memory retains the transfer data input from the other processors in a manner that makes it possible to distinguish between the retained data and the transfer data.
Citation Information
Patent Citations
Data replication for accelerator
US11500802B1
SIMD processor array system and data transfer method thereof
WO2009110497A1