Arithmetic unit

The computing unit optimizes data transmission in systolic arrays by combining normal and reconfiguration data on a shared data bus, reducing reconfiguration time and enhancing efficiency through a time slot method and reception processing units, thus improving calculation speed.

JP2026010559APending Publication Date: 2026-01-22FUJITSU LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2024110510
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-09
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Conventional systolic arrays suffer from inefficient use of data buses and configuration paths, leading to increased reconfiguration times as the scale of the array expands, which hampers overall calculation efficiency.

Method used

The proposed computing unit employs a data bus that transmits both normal calculation data and reconfiguration data, using a time slot method to group transmission data for multiple processing elements (PEs) on the same bus, allowing simultaneous transmission of data for different PEs in separate cycles, and each PE has a reception processing unit to extract the appropriate data.

Benefits of technology

This approach reduces the time required for reconfiguration, optimizes data transmission efficiency, and enhances processing performance by eliminating the need for a dedicated configuration path, thereby improving overall calculation speed.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026010559000001_ABST
    Figure 2026010559000001_ABST
Patent Text Reader

Abstract

To shorten a time required for reconfiguration.SOLUTION: Transmission data to two or more arithmetic element 3 on the same data bus 4 include first data and second data, are collected in one cycle for each transmission data to the same arithmetic element 3 as a transmission destination, are transmitted on the data bus 4 in cycles different for each arithmetic element 3, and each arithmetic element 3 has a reception processing part for selecting and receiving the first data and the second data as transmission destinations by extracting the transmission data from the data bus 4 in a cycle corresponding to the arithmetic element 3 mounted with itself.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a computing unit. [Background technology]

[0002] For example, matrix multiplication calculations are frequently used in DNNs (Deep Neural Networks). Therefore, systems equipped with a computing unit (accelerator) that accelerates matrix multiplication calculations are known as hardware configurations. A known accelerator is a matrix computing unit with a two-dimensional systolic array configuration. In a systolic array configuration, multiple PEs (Processing Elements) are arranged in a two-dimensional lattice pattern.

[0003] Also known is a programmable logic device that includes a plurality of programmable circuits and allows blocks of these programmable circuits to be connected by programmable connection means (see, for example, Patent Documents 1 to 3).Furthermore, it is also known to configure a systolic array using programmable circuits.

[0004] FIG. 12 is a diagram illustrating the configuration of a conventional systolic array.

[0005] In the systolic array shown in Figure 12, multiple PEs arranged in a two-dimensional lattice are connected by a data bus DB. The data bus DB includes buses that cascade-connect the PEs in the row direction and buses that cascade-connect the PEs in the column direction. The data bus DB transmits data used by each PE to perform calculations (normal calculations). In Figure 12, the arrows on the data bus DB indicate the flow of data.

[0006] Each PE receives data (operation results) from the adjacent upstream PE via a data bus DB, performs the operation, and passes the operation results to the adjacent downstream PE via the data bus DB.

[0007] In the conventional systolic array configuration shown in FIG. 12, all PEs are cascade-connected via a single configuration path CP. The configuration path CP transmits configuration information for setting the circuit configuration of each PE, which is a programmable circuit. Each PE reconfigures its circuit based on the configuration information received via this configuration path CP, thereby switching the operation (computation content) of each PE as appropriate. In some cases, each PE performs a different operation (computation).

[0008] In general, when comparing the data bus DB and the configuration path CP, the data bus DB is thicker than the configuration path CP. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] US Patent Application Publication No. 2023 / 0014412 [Patent Document 2] Japanese Patent Publication No. 1-080128 [Patent Document 3] Special Publication No. 10-505993 Summary of the Invention [Problem to be solved by the invention]

[0010] Fig. 13 shows the usage of data buses and configuration paths in a conventional systolic array. Fig. 13 shows an example in which calculations and circuit reconfigurations are repeated, with calculation times and reconfiguration times occurring alternately.

[0011] In Figure 13, the area above the reference line H on the paper indicates the usage status of the data bus DB, and the area below the reference line H on the paper indicates the usage status of the configuration path CP. The right direction on the paper indicates elapsed time. The area indicated by a dot pattern indicates the state in which the data bus DB is used to transmit data, and the data bus DB is used during calculation time. The area indicated by diagonal lines indicates the state in which the configuration path CP is used to transmit configuration information, and the configuration path CP is used during reconfiguration time.

[0012] As shown in FIG. 13, in the conventional systolic array, the configuration path CP is unused during the calculation time, and the data bus DB is unused during the reconfiguration time, resulting in inefficient use.

[0013] In a systolic array, the total time required for calculation can be expressed as the sum of the calculation time and the reconfiguration time. When the scale of the systolic array is expanded by increasing the number of PEs constituting the systolic array, the calculation time is reduced, but the reconfiguration time increases. Therefore, in conventional systolic arrays, there is a need to efficiently transmit configuration information, especially during reconfiguration, in order to reduce the time required for reconfiguration.

[0014] In one aspect, the present invention aims to enable the time required for reconstruction to be reduced. [Means for solving the problem]

[0015] Therefore, this arithmetic unit is an arithmetic unit having an array of arithmetic unit elements configured as programmable circuits whose logical functions can be programmed, and the transmission data transmitted by a data bus connecting two or more of the plurality of arithmetic unit elements includes first data used in one of the two or more arithmetic unit elements and second data used to reconfigure the logical function of one of the two or more arithmetic unit elements, the transmission data for two or more arithmetic unit elements on the same data bus is combined into one cycle for each transmission data having the same arithmetic unit element as the destination and transmitted on the data bus in a different cycle for each arithmetic unit element, and each of the arithmetic unit elements has a receiving processing unit that extracts the transmission data from the data bus in the cycle corresponding to the arithmetic unit element in which it is installed, thereby selecting and receiving the first data and the second data having the same destination as the arithmetic unit element. [Effects of the Invention]

[0016] According to one embodiment, the time required for reconstruction can be reduced. [Brief explanation of the drawings]

[0017] [Figure 1] FIG. 2 is a diagram illustrating a configuration of a group of PEs in an accelerator according to an example of the first embodiment; [Figure 2] FIG. 2 is a diagram for explaining a transmission method of transmission data in an accelerator as an example of the first embodiment. [Figure 3] FIG. 2 is a diagram illustrating an example of a format of transmission data transferred in an accelerator as an example of the first embodiment; [Figure 4] FIG. 2 is a diagram illustrating transmission of transmission data over a plurality of data buses in the accelerator as an example of the first embodiment; [Figure 5] 2 is a diagram illustrating a configuration example of a reception processing unit provided in each PE of the accelerator as an example of the first embodiment; FIG. [Figure 6]10A and 10B are diagrams for explaining the effect of a transmission method of transmission data in an accelerator as an example of the first embodiment. [Figure 7] FIG. 10 is a diagram illustrating a transmission method of transmission data in an accelerator as an example of the second embodiment. [Figure 8] FIG. 10 is a diagram illustrating an example of a data configuration for realizing a data set to be transferred in an accelerator as an example of the second embodiment. [Figure 9] FIG. 10 is a diagram illustrating transmission of transmission data over a plurality of data buses in an accelerator as an example of the second embodiment. [Figure 10] FIG. 10 is a diagram illustrating a configuration of a reception processing unit provided in each PE of an accelerator as an example of a second embodiment. [Figure 11] 10A and 10B are diagrams for explaining the effect of a transmission method of transmission data in an accelerator as an example of the second embodiment. [Figure 12] FIG. 1 illustrates a configuration of a conventional systolic array. [Figure 13] FIG. 1 is a diagram illustrating the usage of a data bus and a configuration bus in a conventional systolic array. DETAILED DESCRIPTION OF THE INVENTION

[0018] Hereinafter, an embodiment of the present computing unit will be described with reference to the drawings. However, the embodiment shown below is merely an example, and it is not intended to exclude various modifications and application of techniques not explicitly stated in the embodiment. In other words, the present embodiment can be implemented with various modifications within the scope of its purpose. Furthermore, each figure does not intend to include only the components shown in the figure, but can also include other functions, etc.

[0019] (I) Description of the First Embodiment (A) Configuration FIG. 1 is a diagram illustrating a configuration of a PE group 6 of an accelerator 1 as an example of the first embodiment.

[0020] The accelerator 1 is a hardware accelerator having a function of performing calculations, and is, for example, a computing unit connected to a host computer (not shown). The accelerator 1 may perform, for example, matrix calculations as an example of calculations, or may perform calculations other than matrix multiplication.

[0021] The host computer may be, for example, an HPC (High Performance Computing) computer, or may be a personal computer, and may be implemented in various modified forms.

[0022] The host computer issues a command to the accelerator 1 to instruct it to execute a calculation, causing the accelerator 1 to perform the calculation. The host computer receives the calculation result from the accelerator 1.

[0023] The host computer transmits data (which may be called data for normal calculation) to be used for executing the calculation (normal calculation) together with a command to execute the calculation to the accelerator 1. The data for normal calculation is an example of first data used in any of two or more PEs 3 (operation unit elements).

[0024] Furthermore, the host computer may cause the accelerator 1 to perform circuit reconfiguration as necessary. For example, the host computer may transmit a command to the accelerator 1 to cause the accelerator 1 to perform circuit reconfiguration. Along with this command, the host computer may transmit information (which may be referred to as reconfiguration data) for setting the circuit configuration of each PE3, which is a programmable circuit, in order to perform circuit reconfiguration of the accelerator 1. The reconfiguration data is an example of second data used to reconfigure the logical function of any of two or more PE3 (operational unit elements).

[0025] As shown in FIG. 1, the accelerator 1 includes a controller 7 and a calculation unit 2.

[0026] The controller 7 controls the execution of matrix multiplication operations by the arithmetic unit 2. The controller 7 receives commands such as commands to execute operations from the host computer and controls the operation of the arithmetic unit 2. Furthermore, when the controller 7 receives a command to perform circuit reconfiguration from the host computer, it transmits information (reconfiguration data) to the target PE3 in the PE group 6 to set the circuit configuration of the PE3, which is a programmable circuit.

[0027] The arithmetic unit 2 performs calculations in accordance with commands issued from the host computer. The arithmetic unit 2 has a two-dimensional systolic array configuration. The arithmetic unit 2 may perform, for example, matrix multiplication calculations. The arithmetic unit 2 includes a PE group 6 having a plurality of PEs 3. The arithmetic unit 2 may have one or more memories that store data before being input to the PE group 6 and calculation results by the PE group 6. The PE3 is an example of a calculation unit element configured as a programmable circuit whose logic functions can be programmed. The circuit configuration of the PE3 is set based on reconfiguration data.

[0028] 1 is a computing unit having a systolic array configuration, and a plurality of PEs 3 are arranged in a two-dimensional lattice in a PE group 6. In the example shown in Fig. 1, the accelerator 1 has a two-dimensional systolic array configuration of 4 rows x 4 columns, consisting of 16 PEs 3.

[0029] In the systolic array (PE group 6) shown in Fig. 1, a plurality of PEs 3 arranged in a two-dimensional lattice are connected by data buses 4. The data buses 4 include a plurality of buses (four in the example of Fig. 1) that cascade connect the PEs 3 in the row direction (horizontal direction in Fig. 1) and a plurality of buses (four in the example of Fig. 1) that cascade connect the PEs 3 in the column direction (vertical direction in Fig. 1).

[0030] The data bus 4 transmits normal calculation data used by each PE 3 to perform calculations (normal calculations). Each PE 3 receives data (calculation results) from an adjacent upstream PE 3 via the data bus 4 and performs the calculation. The results of the calculations executed in the PE 3 (calculation results) are passed to the adjacent downstream PE 3 via the data bus 4.

[0031] Furthermore, in this accelerator 1, the data bus 4 also transmits reconfiguration data for setting the circuit configuration of each PE 3, which is a programmable circuit. Reconfiguration data is transmitted to the data bus 4, with each of the multiple PEs 3 arranged on the same data bus 4 as its destination.

[0032] In this way, in this accelerator 1, the data bus 4 not only functions to transmit normal calculation data but also functions as a configuration path, which is provided in conventional systolic arrays. The data bus 4 is shared by transmitting normal calculation data and reconfiguration data.

[0033] Hereinafter, when there is no particular distinction between the normal calculation data and the reconstruction data transmitted on the data bus 4, they will be referred to as transmission data.

[0034] The control for transmitting the transmission data (normal calculation data, reconstruction data) to the data bus 4 may be performed, for example, by the controller 7, or by another control unit not shown, and can be implemented in various modifications.

[0035] In FIG. 1, the arrows shown on the data bus 4 indicate the flow of transmission data (normal calculation data, reconstruction data). In each PE3, the circuit is reconfigured based on the reconfiguration data received via the data bus 4, and the operation (calculation contents) of each PE3 is switched appropriately. In some cases, each PE3 performs a different operation (calculation).

[0036] FIG. 2 is a diagram for explaining a transmission method of transmission data in the accelerator 1 as an example of the first embodiment.

[0037] In FIG. 2, symbol A indicates a plurality of PEs 3 constituting one of the plurality of columns included in the PE group 6. In FIG. 2, an example is shown in which transmission data is sent to eight PEs 3, PE#1 to PE8, connected to the bus 4. In the expression PE#1 to #8, #1 to #8 are identification information that identify the PEs 3. This identification information may be called a PE number. Furthermore, symbol B illustrates transmission data transmitted to the data bus 4 indicated by symbol A.

[0038] In the accelerator 1 of the first embodiment, transmission data is transmitted to the data bus 4 using a time slot method for each PE3, and transmission data for multiple PE3s on the same data bus 4 is grouped by destination (each PE3), and transmission data addressed to the same PE3 is transmitted in one cycle.

[0039] That is, transmission data for two or more PEs 3 on the same data bus 4 is collected into one cycle for each transmission data destined for the same PE 3, and transmitted on the data bus 4 in a different cycle for each PE 3.

[0040] In the example shown in FIG. 2, two pieces of transmission data for PE#1 are transmitted in cycle #1, and one piece of transmission data for PE#2 is transmitted in cycle #2. Three pieces of transmission data for PE#3 are transmitted in cycle #3, and four pieces of transmission data for PE#4 are transmitted in cycle #4. Furthermore, two pieces of transmission data for PE#6 are transmitted in cycle #6, and one piece of transmission data for PE#7 is transmitted in cycle #7. Furthermore, three pieces of transmission data for PE#8 are transmitted in cycle #8. Note that there is no transmission data for PE#5, so no transmission data is transmitted in cycle #5.

[0041] When the data set shown by symbol B in FIG. 2 is transmitted over the data bus 4, the bus occupation time is 8 cycles.

[0042] In the accelerator 1 of the first embodiment, the procedures required for data transmission (such as destination confirmation and negotiation) only need to be performed in each cycle, and do not need to be performed for each piece of transmission data. Therefore, in the example shown in Fig. 2, 16 pieces of transmission data to PE#1 to #8 are transmitted in an 8-cycle transmission procedure.

[0043] Similar processing is performed for each column, and similar processing is also performed for each of the multiple rows included in the PE group 6.

[0044] As described above, in the accelerator 1 of the first embodiment, transmission data addressed to each of the plurality of PEs 3 arranged on the data bus 4 is transmitted to the data bus 4. As a result, a plurality of transmission data addressed to different PEs 3 are transmitted to one data bus 4.

[0045] Therefore, in the accelerator 1 of the first embodiment, each PE 3 has a reception processing unit 5a (see FIG. 5) for receiving, from the data bus 4, transmission data that the PE 3 processes.

[0046] FIG. 3 is a diagram illustrating an example of a format of transmission data transferred in the accelerator 1 according to the first embodiment.

[0047] The transmission data has a format as a data storage area. In the example shown in Fig. 3, eight data storage areas (transmission data) are provided corresponding to PEs #1 to #8, respectively. A data section is stored in each data storage area. The data section may be, for example, 64-bit data from bits 0 to 63. In the example shown in Fig. 3, the data section is represented as DATA[63:00].

[0048] These data sections are data to be processed by PE3. The data section included in the data for normal calculation is used for calculation by PE3. The data section included in the data for reconstruction is used for reconstruction of PE3. In the transmission data illustrated in FIG. 3, the data section can be said to have a bus width of 64 bits.

[0049] When the amount of information required by the PE 3 matches the bus width of the data bus 4, the transmission data can be transmitted most efficiently.

[0050] FIG. 4 is a diagram illustrating transmission of transmission data on a plurality of data buses 4 in the accelerator 1 as an example of the first embodiment.

[0051] 4 shows multiple pieces of transmission data transmitted on each of four adjacent data buses #00 to #03 in accelerator 1. In addition, the right direction of the paper in FIG. 4 indicates the direction of time passage, and one square corresponds to one cycle (transmission cycle). In FIG. 4, multiple pieces of transmission data lined up vertically on the paper are transmitted in the same cycle.

[0052] In each of the multiple PEs 3 on data buses #00 to #03, the data portion of the normal calculation data is used for calculation processing. The calculation results in each PE 3 are transferred to other PEs 3 on the adjacent data bus 4 in each cycle, and are processed sequentially. For example, the calculation results by a specific PE 3 on data bus #00 are processed in another PE 3 adjacent to this specific PE 3 on data bus #01 in the next cycle.

[0053] In the accelerator 1 of the first embodiment, the transmission data includes a calculation / reconstruction selection signal, a VLD signal, and a data portion.

[0054] The data section is data processed by PE3. The data section may be, for example, 64-bit data from bit 0 to bit 63. In the example shown in FIG. 4, the data section of the transmission data transmitted via data bus #00 is represented as DATA0. The data section of the transmission data transmitted via data bus #01 is represented as DATA1, the data section of the transmission data transmitted via data bus #02 is represented as DATA2, and the data section of the transmission data transmitted via data bus #03 is represented as DATA3.

[0055] In the example shown in FIG. 4, for example, data sections indicated by D00 to D03 and C00 to C03 are transmitted via data bus #00, and data sections indicated by D10 to D13 and C10 to C13 are transmitted via data bus #01.

[0056] The calculation / reconstruction selection signal is a signal that indicates whether the transmission data is data for normal calculation (first data) or data for reconstruction (second data). In the example shown in Fig. 4, Config0 to 3 correspond to the calculation / reconstruction selection signal. Transmission data with 0 set in Config0 to 3 is data for normal calculation (see symbols P1 and P3), and transmission data with 1 set in Config0 to 3 is data for reconstruction (see symbol P2).

[0057] That is, 0 in Config0 to 3, i.e., 0 set in the calculation / reconstruction selection signal, is an example of identification information indicating that the transmission data is data for normal calculation. Also, 1 in Config0 to 3, i.e., 1 set in the calculation / reconstruction selection signal, is an example of identification information indicating that the transmission data is data for reconstruction.

[0058] The VLD signal is a signal that indicates whether the transmission data is valid or not. Valid data may be data that PE3 should receive. PE3, for example, acquires and processes the data portion included in the transmission data for which the VLD signal is set to 1.

[0059] 4, the VLD signal for the transmission data transmitted over data bus #00 is represented as VLD0, the VLD signal for the transmission data transmitted over data bus #01 is represented as VLD1, the VLD signal for the transmission data transmitted over data bus #02 is represented as VLD2, and the VLD signal for the transmission data transmitted over data bus #03 is represented as VLD3.

[0060] In the example shown in Fig. 4, the VLD signal of the transmission data including the data portion is set to 1. The transmission data in which the VLD signal is set to 1 and the calculation / reconstruction selection signals (Config0 to 3) are set to 0 (see the hatched area with diagonal lines in Fig. 4) is processed as data for normal calculation in PE3. Also, the transmission data in which the VLD signal is set to 1 and the calculation / reconstruction selection signals (Config0 to 3) are set to 1 (see the hatched area with grid pattern in Fig. 4) is processed as data for reconstruction in PE3.

[0061] 4, it takes four cycles to transmit the data for reconfiguration on each of the four data buses 4 (see symbol P2). Therefore, strictly speaking, the time required for reconfiguration of PE3 is (4 + α) cycles, which is the sum of these four cycles and the time required for synchronization between the data buses 4.

[0062] FIG. 5 is a diagram illustrating a configuration of a reception processing unit 5a provided in each PE 3 of the accelerator 1 as an example of the first embodiment.

[0063] As shown in FIG. 5, the reception processing unit 5a includes a counter 51 and a determiner 52.

[0064] As described above, in the accelerator 1 of the first embodiment, transmission data is transmitted in a time slot system, and cycles are associated with PEs 3.

[0065] The counter 51 counts the number of cycles. The determiner 52 detects a specific cycle corresponding to the PE 3 in which it is installed (which may be called the own PE 3), and takes in transmission data from the data bus 4 in this specific cycle.

[0066] The determiner 52 may synchronize with a timing adjustment block (not shown) that adjusts the timing of inputting transmission data to the data bus 4, thereby identifying the reception timing (cycle) of transmission data that will be subsequently transmitted to the PE 3 itself.

[0067] In the example shown in Figure 2 above, for example, the receiving processing unit 5a of PE#1 can obtain two pieces of transmission data addressed to its own PE3 (PE#1) by taking in transmission data from the data bus 4 in cycle #1.

[0068] Each receiving processing unit 5a of the multiple PE3s extracts transmission data from the data bus 4 in a cycle associated with the PE3 in which it is installed (its own PE3), thereby selecting and receiving normal calculation data (first data) and reconstruction data (second data) destined for itself.

[0069] (B) Operation In the accelerator 1 of the first embodiment configured as described above, the controller 7 receives commands and the like sent from the host computer and performs processing. The controller 7 sends normal calculation data to the PE3 via the data bus 4 to cause the PE3 to perform calculations. The controller 7 also sends reconfiguration data to the PE3 via the data bus 4 to change the circuit configuration of the PE3.

[0070] The data bus 4 transmits transmission data in a time slot system, and the transmission data is grouped for each PE 3, and transmission data addressed to the same PE 3 is transmitted in one cycle.

[0071] In each PE3, the reception processing unit 5a counts the cycles using the counter 51 and inputs the count result to the decision unit 52. The decision unit 52 takes in the transmission data from the data bus 4 in the cycle corresponding to the PE3 itself. As a result, each PE3 receives the transmission data sent to itself, performs calculations using the received normal calculation data, and also reconfigures the circuit using the received reconfiguration data.

[0072] (C) Effects As described above, according to the accelerator 1 as an example of the first embodiment, the data bus 4 is used to transmit the reconfiguration data to the PE 3, eliminating the need for a dedicated path (config path) for transmitting the reconfiguration data. This reduces the cost of wiring and the like, and also reduces the mounting area.

[0073] Furthermore, multiple pieces of transmission data addressed to multiple PEs 3 are transmitted in a time slot system on the data bus 4, and transmission data with the same destination (PE 3) are transmitted together in one cycle. This reduces the time required for transmission compared to the conventional method of transmitting transmission data one by one through P2P (Peer-to-Peer) communication. Furthermore, the data bus 4 can be used efficiently, improving data transmission efficiency. This also improves the processing performance of the accelerator 1.

[0074] By reducing the time required to transmit the data for reconfiguration to each PE3, the time required to reconfigure the PE3 can be reduced.

[0075] FIG. 6 is a diagram for explaining the effect of the transmission method of the transmission data in the accelerator 1 as an example of the first embodiment.

[0076] In FIG. 6, symbol A indicates a configuration example of the PE group 6, and symbol B indicates a transmission method of transmission data in the accelerator 1 of the first embodiment.

[0077] For example, in a systolic array in which 64 PEs 3 are formed in an 8x8 matrix configuration as shown by symbol A, the conventional method requires one cycle to transmit transmission data to one PE. Therefore, in a systolic array with 64 PEs, it requires 64 cycles to transmit transmission data to all PEs.

[0078] In contrast to this, in the accelerator 1 of the first embodiment, as shown by symbol B in FIG. 6, it is possible to transmit transmission data to all PEs 3 in eight cycles, thereby reducing the time required to transmit transmission data.

[0079] (II) Description of the Second Embodiment (A) Configuration FIG. 7 is a diagram for explaining a transmission method of transmission data in the accelerator 1 as an example of the second embodiment.

[0080] In the second embodiment, as in the first embodiment, the accelerator 1 is a computing unit having a systolic array configuration. PE3 is an example of a computing unit element configured as a programmable circuit whose logic functions can be programmed. The circuit configuration of PE3 is set based on reconfiguration data.

[0081] The accelerator 1 of the second embodiment may have a configuration similar to that of the accelerator 1 of the first embodiment illustrated in FIG.

[0082] In FIG. 7, symbol A indicates a plurality of PEs 3 constituting one of the plurality of columns included in the PE group 6. As in FIG. 2, FIG. 7 also shows an example in which transmission data is sent to eight PEs 3, PE#1 to PE#8, connected to the bus 4. In the expression PE#1 to #8, #1 to #8 are identification information that identify the PEs 3. This identification information may be called a PE number. Furthermore, symbol B illustrates transmission data transmitted to the data bus 4 indicated by symbol A.

[0083] In the accelerator 1 of the second embodiment, transmission data for multiple PEs 3 on the same data bus 4 is collected and formed into one data set, as indicated by the symbol B. Fig. 7 shows an example in which one data set is formed by multiple transmission data for eight PEs 3.

[0084] The example indicated by symbol B in FIG. 7 shows a data set in which multiple pieces of transmission data to be transmitted to PEs #1 to #8 are grouped together (blocked) as one block. In this data set, multiple pieces of transmission data are arranged in ascending order of PE number, with the bottom left at the top and the top right at the bottom, with the rows going leftward and the columns going up. In this case, transmission data with a short required data length is arranged closely together. In addition, if there is a PE3 that does not need to transmit data, the addition of transmission data to that PE3 is skipped, and the transmission data to the next PE3 is incorporated into the data set. In the example shown in FIG. 7, there is no transmission data to PE #5.

[0085] In the accelerator 1 of the second embodiment, transmission data for two or more PEs 3 (operating unit elements) on the same data bus 4 is collected into one data set and transmitted over the data bus 4.

[0086] When transmitting the data set illustrated by symbol B in FIG. 7 over the data bus 4, one row of transmission data may be transmitted in one cycle (transmission cycle). For example, in cycle #1, two transmission data for PE#1, one transmission data for PE#2, and one transmission data for PE#3 are transmitted. In the next cycle #2, two transmission data for PE#3 and two transmission data for PE#4 are transmitted. In cycle #3, two transmission data for PE#4 and two transmission data for PE#6 are transmitted, and in cycle #4, one transmission data for PE#7 and three transmission data for PE#8 are transmitted.

[0087] Therefore, when the data set illustrated by symbol B in FIG. 7 is transmitted over the data bus 4, the bus occupation time is four cycles.

[0088] In the accelerator 1 of the second embodiment, the transmission procedure (destination confirmation, negotiation, etc.) only needs to be performed once before transmitting a data set, and does not need to be performed for each cycle. Therefore, in the example shown in Fig. 7, 16 pieces of transmission data to PE#1 to #8 are transmitted in one transmission procedure.

[0089] Similar processing is performed for each column, and similar processing is also performed for each of the multiple rows included in the PE group 6.

[0090] In the accelerator 1 of the second embodiment, transmission data addressed to each of the plurality of PEs 3 arranged on the data bus 4 is also transmitted to the data bus 4. As a result, a plurality of transmission data addressed to different PEs 3 is transmitted to one data bus 4. This control may be performed by the controller 7, as in the first embodiment, or by another control unit (not shown), and can be implemented in various modifications.

[0091] Therefore, in the accelerator 1 of the second embodiment, each PE 3 has a reception processing unit 5b (see FIG. 10) for receiving, from the data bus 4, transmission data that the PE 3 processes.

[0092] FIG. 8 is a diagram illustrating an example of a data structure for realizing a data set to be transferred in the accelerator 1 as an example of the second embodiment.

[0093] In FIG. 8, symbol A illustrates a data set of transmission data transmitted to PEs #1 to #8. In the data set illustrated by symbol A, multiple pieces of transmission data are arranged in ascending order of PE number, starting from the bottom left of the page, as described above, and therefore, transmission data with the same destination (PE3) are consecutive. A set of consecutive transmission data with the same destination (PE3) may be referred to as a per-PE transmission data set. The number of transmission data included in a per-PE transmission data set may be referred to as the number of slots. In the accelerator 1 of the second embodiment, such a data set is transmitted in multiple cycles (four cycles in the example illustrated in FIG. 8). Furthermore, in FIG. 8, symbol B indicates a data configuration that realizes the data set illustrated by symbol A.

[0094] In the accelerator 1 of the second embodiment, a plurality of segments are generated by dividing the data bus 4 into predetermined fixed lengths. In the data configuration shown by symbol B in Fig. 8, each rectangle corresponds to a segment, and information on the transmission data is set in these segments.

[0095] In the data configuration illustrated by symbol B, information on a plurality of sets of transmission data for each PE is registered. In the data configuration illustrated by symbol B, each segment corresponds to transmission data.

[0096] When a set of transmission data per PE includes a plurality of pieces of transmission data, the first segment among the plurality of segments corresponding to one set of transmission data per PE may be called the first segment.

[0097] The information of the per-PE transmission data set includes a data length, PE specific information, and a data section. The data length indicates the data length of the per-PE transmission data set. The data length is an example of data length information. The data length may be expressed as the number of slots. In the example shown in FIG. 8, L[1:0] indicates the data length, and the number of slots is expressed as a 2-bit value.

[0098] For example, in the data set indicated by symbol A in Figure 8, the PE-specific transmission data set transmitted to PE#1 includes two (two slots) of transmission data, so the value of L[1:0] is set to 1, which corresponds to 2.

[0099] The PE identification information is information that identifies the PE 3, and is an example of destination identification information. In the example shown in Fig. 8, PE[3:0] indicates the PE identification information, and the PE number is expressed as a 4-bit value.

[0100] For example, in the data set indicated by symbol A in Fig. 8, the PE number of the per-PE transmission data set transmitted to PE#1 is 1, so the value of PE[3:0] is set to 0, which corresponds to 1. The data section is usually data for calculation or data for reconstruction.

[0101] When a set of transmission data for each PE includes multiple pieces of transmission data, among the data length, PE specific information, and data part in multiple segments corresponding to one set of transmission data for each PE, the data length and PE specific information are stored in the first segment, whereas the data part is stored in each segment.

[0102] 8, DATA[9:0] indicates the data portion stored in the first segment, and indicates that the data portion is stored as a 10-bit value. In addition, in segments other than the first segment, DATA[15:0] indicates the data portion, and indicates that the data portion is stored as a 16-bit value.

[0103] In the example shown in Figure 8, the bottom left corner of the blocked data structure is the beginning of the data structure, and the beginning is where the first segment of the PE-specific transmission data set for PE#1 is located (see symbol P1 in Figure 8).

[0104] By referring to the data length L[1:0] stored in this first segment and skipping the number of segments equivalent to this data length (2 in this example) (see symbol P2 in Figure 8), the first segment of the PE-specific transmission data set for PE#2, which follows PE#l, is reached (see symbol P3 in Figure 8).

[0105] In addition, in the PE-by-PE transmission data set of PE#2, the data length L[1:0] stored in the first segment is referenced, and the number of segments equivalent to this data length (1 in this example) is skipped (see symbol P4 in Figure 8), thereby reaching the first segment of the PE-by-PE transmission data set of PE#3, which follows PE#2 (see symbol P5 in Figure 8).

[0106] In addition, in the PE-by-PE transmission data set of PE#3, the data length L[1:0] stored in the first segment is referenced, and the number of segments equivalent to this data length (3 in this example) is skipped (see symbol P6 in Figure 8), thereby reaching the first segment of the PE-by-PE transmission data set of PE#4, which follows PE#3 (see symbol P7 in Figure 8).

[0107] In addition, in the PE-by-PE transmission data set of PE#4, the data length L[1:0] stored in the first segment is referenced, and the number of segments equivalent to this data length (4 in this example) is skipped (see symbol P8 in Figure 8), thereby reaching the first segment of the PE-by-PE transmission data set of PE#5, which follows PE#4 (see symbol P9 in Figure 8).

[0108] In addition, in the PE-by-PE transmission data set of PE#5, by referring to the data length L[1:0] stored in the first segment and skipping the number of segments equivalent to this data length (1 in this example) (see symbol P10 in Figure 8), the first segment of the PE-by-PE transmission data set of PE#6, which follows PE#5, is reached (see symbol P11 in Figure 8).

[0109] In addition, in the PE-by-PE transmission data set of PE#6, by referring to the data length L[1:0] stored in the first segment and skipping the number of segments equivalent to this data length (1 in this example) (see symbol P12 in Figure 8), the first segment of the PE-by-PE transmission data set of PE#7, which follows PE#6, is reached (see symbol P13 in Figure 8).

[0110] In the PE-by-PE transmission data set of PE#7, the data length L[1:0] stored in the first segment is referenced, and the number of segments corresponding to this data length (1 in this example) is skipped (see symbol P14 in Figure 8), thereby reaching the first segment of the PE-by-PE transmission data set of PE#8, which follows PE#7 (see symbol P15 in Figure 8).

[0111] In the data structure configured in this way, by referring to each leading segment, it is possible to extract a set of data transmitted by each PE from the set of blocked data.

[0112] The receiving processing unit 5b selects and receives from each PE3 the normal calculation data (first data) and the reconstruction data (second data) that are destined for the PE3 itself from the data set based on the data length (data length information) and the PE identification information (destination identification information).

[0113] FIG. 9 is a diagram illustrating transmission of transmission data on a plurality of data buses 4 in the accelerator 1 as an example of the second embodiment.

[0114] 9 shows a plurality of transmission data transmitted on each of four adjacent data buses #00 to #03 in accelerator 1. In addition, the right direction of the paper in FIG. 9 indicates the direction of time passage, and one square corresponds to one cycle (transmission cycle). In FIG. 9, a plurality of transmission data lined up vertically on the paper are transmitted in the same cycle.

[0115] In the PE3 on the data buses #00 to #03, the data portion of the normal calculation data is used for calculation processing. The calculation results in each PE3 are transferred to the adjacent PE3 on the data bus 4 in each cycle, and are processed sequentially. For example, the calculation results by a specific PE3 on the data bus #00 are transferred to another PE3 adjacent to this specific PE3 on the data bus #01 in the next cycle, and are processed sequentially.

[0116] In the accelerator 1 of the second embodiment, the transmission data also includes a calculation / reconstruction selection signal, a VLD signal, and a data portion, as in the first embodiment.

[0117] 9, the data portion of the transmission data transmitted via data bus #00 is represented as DATA0. The data portion of the transmission data transmitted via data bus #01 is represented as DATA1, the data portion of the transmission data transmitted via data bus #02 is represented as DATA2, and the data portion of the transmission data transmitted via data bus #03 is represented as DATA3.

[0118] In the example shown in FIG. 9, for example, data sections indicated by D00 to D03 and C00 to C03 are transmitted via data bus #00, and data sections indicated by D10 to D13 and C10 to C13 are transmitted via data bus #01.

[0119] In the example shown in FIG. 9, Config0 to 3 correspond to the calculation / reconstruction selection signal, and transmission data for which Config0 to 3 are set to 0 is data for normal calculation, and transmission data for which the calculation / reconstruction selection signal is set to 1 is data for reconstruction.

[0120] In the example shown in FIG. 9, the transmission of normal calculation data indicated by diagonal hatching and the transmission of reconstruction data indicated by grid-like hatching are mixed on each of the data buses #00 to #03.

[0121] As described above, in the accelerator 1 of the second embodiment, the data set has PE-specific information in the first segment for each PE-transmitted data set, so even if the reconstruction data is transmitted piecemeal (for each PE3), the destination PE3 can receive the reconstruction data that it should receive.

[0122] That is, by transmitting the normal calculation data and the reconstruction data to the plurality of PEs 3 as a data set in which the data is mixed on the data bus 4, the data can be transmitted to each of the plurality of PEs 3.

[0123] FIG. 10 is a diagram illustrating a configuration of a reception processing unit 5b provided in each PE 3 of the accelerator 1 as an example of the second embodiment.

[0124] As shown in FIG. 10, the reception processing unit 5b includes a holding unit 53, a checking unit 54, and an importing unit 55.

[0125] As described above, in the accelerator 1 of the second embodiment, transmission data for multiple PEs 3 on the same data bus 4 is transmitted as one data set, and each leading segment of multiple PE-specific transmission data sets included in this data set includes a data length, PE specific information, and a data portion.

[0126] The storage unit 53 temporarily stores data sets transmitted on the data bus 4. The storage unit 53 may be, for example, a memory. The storage unit 53 may receive data sets that are divided and transmitted over multiple cycles (four cycles in the example shown in FIG. 8), combine the data sets, and store them as a single data set.

[0127] The acquisition unit 55, in accordance with instructions from the check unit 54 (described later), extracts data to be transmitted from the data set stored in the holding unit 53. The PE3 uses the data extracted by the acquisition unit 55 to perform calculations and reconfigure the circuit.

[0128] The check unit 54 refers to the PE identification information of each leading segment of the data set stored in the holding unit 53, and determines whether the PE-specific transmission data set including the leading segment corresponds to the own PE 3.

[0129] For example, the check unit 54 first refers to the PE identification information of the first segment at the beginning of the data set stored in the holding unit 53, and determines whether the per-PE transmission data set corresponds to the own PE 3. If the per-PE transmission data set corresponds to the own PE 3, for example, the check unit 54 notifies the acquisition unit 55 of the storage location of the data to be transmitted that is included in the per-PE transmission data set, and causes the own PE 3 to acquire the data.

[0130] Also, if it does not correspond to the PE3 itself, the check unit 54 refers to the data length L[1:0] of the first segment and skips the number of segments in the data set that correspond to this data length, thereby accessing the first segment of the next PE-specific transmission data set.

[0131] The check unit 54 refers to the PE identification information of this first segment and determines whether the per-PE transmission data set corresponds to its own PE 3. If it does not correspond to its own PE 3, the check unit 54 refers to the data length L[1:0] of the first segment and skips the number of segments equivalent to this data length in the data set, thereby accessing the first segment of the next per-PE transmission data set.

[0132] The checking unit 54 refers to the PE identification information of the first segment of the next per-PE transmission data set accessed in this manner, and determines whether the per-PE transmission data set corresponds to its own PE 3. When the checking unit 54 determines that the per-PE transmission data set corresponds to its own PE 3, it notifies the capturing unit 55 of this fact. At this time, the checking unit 54 may notify the capturing unit 55 of the position of the data portion included in the per-PE transmission data set.

[0133] The acquisition unit 55 extracts the data portion of the per-PE transmission data set corresponding to the PE3 itself from the data set stored in the holding unit 53 in accordance with the instruction of the check unit 54. By repeating the same process thereafter, the reception processing unit 5b acquires transmission data addressed to the PE3 itself from the data bus 4. The PE3 uses the data extracted by the acquisition unit 55 to perform calculations and reconfigure the circuit.

[0134] (B) Operation In the accelerator 1 of the second embodiment configured as described above, the controller 7 receives commands and the like sent from the host computer and performs processing. The controller 7 sends normal calculation data to the PE3 via the data bus 4 to cause the PE3 to perform calculations. The controller 7 also sends reconfiguration data to the PE3 via the data bus 4 to change the circuit configuration of the PE3.

[0135] In the accelerator 1 of the second embodiment, transmission data for multiple PEs 3 on the same data bus 4 is formed as one data set. The data set formed in this way is divided for each cycle and transmitted on the data bus 4. At this time, it is desirable to divide the data set in accordance with the bandwidth of the data bus 4. For example, by dividing the data set so that it has the same size as the bandwidth of the data bus 4, the data bus 4 can be used efficiently.

[0136] In the reception processing unit 5b of each PE3, the holding unit 53 receives a data set that is divided and transmitted over a plurality of cycles, and stores it as one data set.

[0137] Furthermore, the check unit 54 determines whether the per-PE transmission data set stored in the holding unit 53 corresponds to the own PE 3 based on the data length of the first segment of each per-PE transmission data set and the PE identification information.

[0138] Then, the acquisition unit 55 extracts the data portion included in the PE-specific transmission data set that the check unit 54 has determined corresponds to the PE 3. The PE 3 uses the data extracted by the acquisition unit 55 to perform calculations and reconfigure the circuit.

[0139] (C) Effects In the accelerator 1 as an example of the second embodiment, similarly to the first embodiment, the data bus 4 is used to transmit the reconfiguration data to the PE 3, so a dedicated path (config path) for transmitting the reconfiguration data is not required. This makes it possible to reduce the cost of wiring etc. and also to reduce the mounting area.

[0140] Furthermore, transmission data for a plurality of PEs 3 on the same data bus 4 is formed as one data set. The data set formed in this way is divided for each cycle and transmitted to the data bus 4.

[0141] When transmitting a data set, the data set can be divided to fit the bandwidth of the data bus 4, thereby making it possible to use the data bus 4 efficiently and improving the data transmission efficiency. This also makes it possible to improve the processing performance of the accelerator 1.

[0142] The data set includes a plurality of transmission data sets for each PE, and the first segment of each transmission data set for each PE includes the data length, PE identification information, and a data portion as information of the transmission data set for each PE.

[0143] In the reception processing unit 5b of each PE3, the check unit 54 determines whether the per-PE transmission data set corresponds to the own PE3 based on the data length of the first segment of each per-PE transmission data set and the PE identification information in the data set acquired from the data bus 4. This allows each PE3 to acquire the data portion corresponding to each PE3 from the data set that combines the transmission data for multiple PE3s.

[0144] FIG. 11 is a diagram for explaining the effect of the transmission method of the transmission data in the accelerator 1 as an example of the second embodiment.

[0145] In FIG. 11, symbol A indicates a configuration example of the PE group 6, and symbol B indicates a transmission method of transmission data in the accelerator 1 of the second embodiment.

[0146] For example, in a systolic array in which 64 PEs 3 are formed in an 8 × 8 matrix configuration as shown by symbol A, in the conventional method, one cycle is required to transmit transmission data to one PE 3. Therefore, in a systolic array having 64 PEs 3, 64 cycles are required to transmit transmission data to all PEs 3.

[0147] In contrast to this, in the accelerator 1 of the second embodiment, as shown by symbol B in FIG. 11, it is possible to transmit transmission data to all PEs 3 in four cycles, thereby reducing the time required to transmit transmission data.

[0148] Furthermore, by reducing the time required to transmit the data for reconfiguration to PE3, the time required to reconfigure PE3 can be reduced.

[0149] (III) Other The configurations and processes of this embodiment can be selected as needed, or can be combined as appropriate.

[0150] The disclosed technology is not limited to the above-described embodiment, and can be implemented in various modifications without departing from the spirit of the present embodiment.

[0151] Furthermore, the above disclosure will enable those skilled in the art to implement and manufacture the present embodiment. [Explanation of symbols]

[0152] 1. Accelerator 2 Arithmetic section 3 PE 4 Data Bus 5a, 5b Receiving processing section 6 PE group 7 Controller 51 Counter 52 Judgment device 53 Holding part 54 Check Section 55 Intake section

Claims

1. A computing unit in which a plurality of computing unit elements configured as programmable circuits capable of programming logic functions are arranged, transmission data transmitted by a data bus connecting two or more of the plurality of arithmetic unit elements includes first data used in any of the two or more arithmetic unit elements and second data used to reconfigure a logical function of any of the two or more arithmetic unit elements, the transmission data for two or more of the arithmetic unit elements on the same data bus are collected into one cycle for each transmission data having the same arithmetic unit element as a destination, and are transmitted on the data bus in different cycles for each of the arithmetic unit elements; each of the arithmetic unit elements has a receiving processing unit that selects and receives the first data and the second data, each of which is a destination of the arithmetic unit element, by extracting the transmission data from the data bus in a cycle associated with the arithmetic unit element in which the arithmetic unit element is installed; A computing unit characterized by:

2. A computing unit in which a plurality of computing unit elements configured as programmable circuits capable of programming logic functions are arranged, transmission data transmitted by a data bus connecting two or more of the plurality of arithmetic unit elements includes first data used in any of the two or more arithmetic unit elements and second data used to reconfigure a logical function of any of the two or more arithmetic unit elements, the transmission data for two or more of the plurality of arithmetic unit elements on the same data bus is collected to form one data set and transmitted over the data bus; data length information and destination specification information for each set of transmission data having the same arithmetic unit element as a destination, among the data sets; each of the computing unit elements has a receiving processing unit that selects and receives the first data and the second data, each of which is destined for the computing unit element, from the data set based on the data length information and the destination specification information; A computing unit characterized by:

3. The first data has identification information indicating that it is the first data, and the second data has identification information indicating that it is the second data.

3. The computing unit according to claim 1 or 2.

Citation Information

Patent Citations

  • JP1975005993A

  • Programmable logic device

    JP1989080128A

  • Method and system for enhancing programmability of a field-programmable gate array via a dual-mode port

    US20230014412A1