Compilation for a Synchronous Processor

The method of generating a fully scheduled intermediate representation for synchronous processors addresses the complexity of programming errors, enhancing accuracy and flexibility in hardware configurations.

JP7708921B2Active Publication Date: 2025-07-15GOOGLE LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024066659
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-08-22
Filing Date
2024-04-17
Publication Date
2025-07-15
Estimated Expiration
2040-08-21

AI Technical Summary

Technical Problem

Programming synchronous processors is highly error-prone due to the complexity of scheduling operations, leading to inaccurate or undefined calculation results.

Method used

A method for compiling programs for synchronous processors that automatically generates a fully scheduled intermediate representation, reducing human error and enabling reuse across different hardware configurations, while reporting errors if no valid output exists.

Benefits of technology

This approach significantly reduces programming errors and allows programs to be reused across various hardware configurations without rewriting, ensuring accurate and consistent operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007708921000003
    Figure 0007708921000003
  • Figure 0007708921000004
    Figure 0007708921000004
  • Figure 0007708921000005
    Figure 0007708921000005
Patent Text Reader

Abstract

To compile a program free from the influence of latency for a synchronous processor.SOLUTION: A method according to one embodiment of the present invention includes a step of receiving an intermediate expression of a program defining actions to be executed by a plurality of respective constituent elements of a synchronous processor. The intermediate expression assigns respective clock cycle values to be scheduled to each of a plurality of actions such that the actions are executed by the synchronous processor. The intermediate expression is processed to generate respective update windows for actions in the intermediate expression to request a hardware configuration update, and the update window defines a time range in which a configuration update instruction can be executed to perform the hardware configuration update. The configuration update instruction is scheduled to occur during one or a plurality of update windows according to configurational constraint of the synchronous processor.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to compiler techniques for synchronous integrated circuit accelerators.

Background Art

[0002] A synchronous integrated circuit accelerator is an application - specific integrated circuit (ASIC) designed to perform highly parallel synchronous operations. This parallelism is achieved by integrating many different independent processing elements that can be executed simultaneously. Such devices are well - suited for accelerating inference paths through neural networks. For example, each of the independent processing elements performs a different multiplication or addition of the layer input using weights.

[0003] For the sake of brevity, in this specification, such devices are referred to as synchronous processors or synchronous chips. In practice, such devices can be implemented using multiple silicon wafers, for example, using die stacking or silicon interposers.

[0004] In this specification, when a processor is synchronous, it means that the operations performed by independent processing units do not perform branched execution, such as an if / else statement in an instruction program. Rather, the operations can be scheduled either partially or fully in advance. For example, the operations of some synchronous processors must be scheduled at the individual cycle level, meaning that every operation of every processing element is assigned to a specific slot within the execution cycle.

Summary of the Invention

Problems to be Solved by the Invention

[0005] Due to this complexity, programming a synchronous processor is highly error-prone. For example, a user bug that causes an unexpected delay in a single cycle can result in inaccurate or undefined calculation results, which can mean that multiple paths of the same program can produce different results.

Means for Solving the Problem

[0006] This specification describes techniques for compiling a program written for a synchronous processor. The techniques described below relate to scheduling data instructions that operate on data flowing through the processor. The techniques described below also relate to automatically generating and scheduling configuration update instructions executed by the control hardware of the processor.

[0007] Certain embodiments of the subject matter described in this specification may be implemented to realize one or more of the following advantages. The system can automatically generate a fully scheduled intermediate representation for a synchronous processor from an input program that is not affected by latency. This greatly reduces human error when programming a synchronous processor. In addition, a program that is not affected by latency can be reused for different hardware configurations without being rewritten. The intermediate representation enables configuration update instructions to be fully scheduled and gives the compiler the ability to report errors when no output program may exist.

[0008] Details of one or more embodiments of the subject matter of this specification are set forth in the following accompanying drawings and description. Other features, aspects, and advantages of the subject matter will become apparent from the description, drawings, and claims.

Brief Description of the Drawings

[0009]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

[0010] Like reference numbers and designations in the various drawings indicate like elements.

[0011] FIG. 1 is a flowchart of an exemplary process for generating an intermediate representation of an input program. For convenience, the process is described as being executed by a system of one or more computers, located at one or more locations, and appropriately programmed according to this specification.

[0012] The system receives (110) a program that defines the operations to be performed by the components of a synchronous processor. Generally, each statement of a program can define one or more input variables, output variables, and operations. The program need not define latency or timing information as to when the operations are to be executed. Instead, the system can then use variable names to determine the timing constraints for executing the program on the synchronous processor.

[0013] TABLE 1 lists an exemplary program that multiplies two values read from RAM, adds the result to a third value read from RAM, and writes the result to RAM.

[0014] **[Table 1]**

[0015] Arrow syntax, such as ("<-name-<"), identifies subprograms that are on-chip parameters supplied from the right and assigns those on-chip results to the name on the left. In the case of a synchronous processor, the arrow syntax can be used to refer to specific hardware components on the processor. In this example, "addr" is used as an input to the overall composite program as declared by "proc addr->..." on line 1. In this case, "addr" specifies which RAM of the processor components should be read from in each of the three read operations. For example, the RAM represented by "addr" can calculate and store the partial sums that should be used by the next stage of the program.

[0016] In particular, the exemplary program is not affected by latency, which means that it does not specify any information about the latency of the components and does not specify the timing required to execute each operation. This means that, unlike conventional programs, the program in TABLE 1 (Table 1) does not require statements to be executed sequentially. In practice, statements are often expected to be executed in different orders or in parallel, depending on the nature of the synchronous processor.

[0017] The system obtains (120) the respective clock latency values for each of the components of the synchronous processor. The system can maintain a mapping, such as a table, between (1) each operation recognized by a keyword or function in the program and (2) the respective clock latency values representing the number of clock cycles required for each operation on the synchronous processor. This configuration enables the same latency-unaffected program to be compiled for various basic hardware designs. This is advantageous because the same program can be compiled for the new design of the hardware when the design of the basic hardware changes some of the component latency values.

[0018] The system generates clock timing constraints (130) for executing a program. The clock timing constraints define when a particular operation must occur relative to other operations. In general, a value read must occur at least one clock cycle before the value is used. For example, a multiplication operation that requires two inputs has a constraint that the two inputs must be read at least one clock cycle before the multiplication operation begins. Thus, for example, the multiplication operation in row 5 of TABLE 1 must start at least one clock cycle after each of the read operations in rows 3 and 4. In addition, consecutive operations should not overlap in time. The system can generate constraints by using variable names in the program to determine which operations must occur after which other operations.

[0019] FIG. 2 is a graphical representation 200 of an exemplary program. The graphical representation 200 corresponds to the exemplary program listed in TABLE 1. The graphical representation 200 has nodes each corresponding to one of the operations in the program. Thus, the graph has a read node 205 representing the read command in row 2, a read node 210 representing the read command in row 3, a read node 215 representing the read command in row 4, a multiplication node 220 representing the multiplication command in row 5, an addition node 225 representing the addition command in row 6, and a write node 230 representing the write command in row 7.

[0020] To generate clock timing constraints, the system can scan the program, or equivalently the edges in the graph representation, to determine dependencies based on variable names. The system can then calculate the required clock latency values for the discovered dependencies between operations. For example, if the multiplication operation on line 4 takes 4 cycles, the system can generate a constraint that the addition operation on line 5 depends on the output of the multiplication operation and should not start within 4 cycles after the multiplication operation starts.

[0021] As another example, the addition operation on line 6 depends on variable x, which is read from line 2. Thus, there is a dependency between the operation on line 6 and the operation on line 2.

[0022] As shown in FIG. 2, the two read nodes 205 and 210 have a latency of 1 cycle. The read node 215 has a latency of 2 cycles. The multiplication node 220 has a latency of 4 cycles, the addition node 225 has a latency of 3 cycles, and the write node 230 has a latency of 1 cycle.

[0023] As shown in FIG. 1, the system generates (140) an intermediate representation of the program that assigns respective clock cycle values to each of a plurality of operations. The intermediate representation is intermediate in the sense that it is lower-level code than the user program but generally not so low-level that it can be directly executed by a synchronous processor.

[0024] The system can generate the intermediate representation by starting with any arbitrary operation in the program and then scan for dependencies and latencies. For example, the system can start at the read node 210 and assign its read operation as being executed at an arbitrary start cycle number, for example cycle 0.

[0025] The system can then assign the subsequent multiplication operation as being executed in cycle 1, using the dependency between node 210 and node 220, as well as the 1-cycle latency of node 210.

[0026] The system can then investigate other dependencies of the multiplication node 220, which is the read node 215. The read node 215 represents a read from a scaling RAM that has different characteristics than the RAM used for the read node 210. Therefore, the latency is different and is 2 cycles instead of 1 cycle. Thus, the system assigns the read operation of the read node 215 as being executed 2 cycles before the multiplication operation of the multiplication operation 220.

[0027] This results in a negative cycle value of -1. However, the values in the intermediate representation represent only the relative cycle times, and thus negative values are allowed, and in some cases expected, or even inevitable.

[0028] The system can then investigate the dependencies of the addition node 225 corresponding to the addition operation. Since the latency between the multiplication node 220 and the addition node 225 is 4 cycles, the system assigns the addition operation as being executed 4 cycles after the multiplication operation, for example in cycle 5.

[0029] The system can then investigate other dependencies on the addition node 225, which is the read node 205. Due to the 1-cycle latency of the read node 205, the system can assign the read operation as being executed 1 cycle before the addition node 225. Thus, the system assigns a cycle value of 4 to the read operation.

[0030] Finally, the system examines the dependencies of the addition node 225, which in this example is the write node 230 corresponding to the write operation. Due to the 3-cycle latency of the addition node 225, the system assigns a cycle value of 8 to the write operation.

[0031] Accordingly, the intermediate representation associates each cycle time value with each operation in the program. The system can store the intermediate representation using any suitable data structure. For example, the system can store a mapping between (1) the cycle time value and (2) the instruction identifier in a node in the graph, the instruction text, or some other representation of the operation.

[0032] An exemplary intermediate representation resulting from processing the graph shown in the values of FIG. 2 is shown in TABLE 2.

[0033]

Table 2

[0034] Unlike the original input program, the statements in TABLE 2 do not define the chronological order of the operations to be executed by the synchronous processor. This order of operations will very likely succeed in the actual hardware because all latency information for scheduling is automatically generated at the cycle level.

[0035] In some cases, the correct intermediate representation does not exist for the input program. In such cases, the system can issue an error indicating that it cannot compile the program correctly. This can occur, for example, when the dependencies in the input program form a cycle. This can also occur when the hardware cannot generate a value early enough for it to be available when needed.

[0036] As another example, any kind of operation having multiple inputs derived from the same intermediate value may be incorrectly formed. If the value "a" is generated in cycle 0, the value "b" is calculated from "a" in 3 cycles, the value "c" is calculated from the value "a" in 2 cycles, and the program attempts to multiply "b" and "c" without dealing with the latency difference, there is no correct way to schedule those operations. Instead, the user may not need to add code to communicate "c" to the next cycle, for example, by using RAM or a delay register.

[0037] The intermediate representation defines a cycle-level schedule of operations to be executed by a synchronous processor, but the compiler can issue additional instructions to actually execute the program on the components of the synchronous processor. Specifically, the compiler may need to generate a configuration update instruction that defines a change in the data path between the processing components of the synchronous processor. This is particularly true when the instructions have polyhedral loop semantics, meaning that each sequence of instructions may be executed multiple times. Thus, for example, the program semantics can define that a particular read operation should be read from a particular register in 10 consecutive iterations, and then should be read from a different particular register in the next 10 iterations.

[0038] To switch from one configuration to another, the processor can issue a configuration update instruction that changes the inputs to the multiplexers in the data path between the processor components.

[0039] Figure 3 is a flowchart of an exemplary process for issuing a configuration update instruction for an intermediate representation. For convenience, the process is described as being executed by one or more computer systems, located in one or more locations, and appropriately programmed in accordance with this specification.

[0040] The system processes the intermediate representation to generate respective update windows for each operation (310). The system processes the instructions in the intermediate representation and generates respective update windows that represent the time ranges during which the corresponding configuration update instructions need to be executed for each instruction that requests a configuration update. In some implementations, in the absence of received configuration update instructions, the processor uses the same configuration when the next operation is executed.

[0041] The system obtains one or more configuration constraints for the synchronous processor (320). One of the natural configuration constraints for the control hardware of the synchronous processor is how many configuration update instructions can be executed in a single cycle. For example, some synchronous processors may support executing only one, two, four, or eight configuration update instructions simultaneously.

[0042] Other configuration constraints may include in which cycles configuration instructions can be executed. For example, in some implementations, configuration update instructions can be executed only every N cycles, where N is 2, 4, or 8, for example.

[0043] Other configuration constraints may include power region constraints that define how many configuration updates can be executed per power region of the processor.

[0044] The system generates and schedules configuration update commands from update windows and configuration constraints (330). Each configuration update command can change how the command behaves for different loop iterations. For example, a configuration update command can change the register for a register read operation. As another example, a configuration update command can change the arguments of a multiplexer in the data path of a synchronous processor. This can cause downstream operations to use different arguments in subsequent loop iterations. In general, each configuration update command is scheduled to occur within one of the update windows generated from the processing of the intermediate representation. The system can use any suitable constraint resolution procedure to schedule all the required configuration update commands within the available update windows. An exemplary constraint resolution example is described here.

[0045] Figure 4 is a diagram showing the scheduling of configuration update commands under the constraints of a synchronous processor. Figure 4 has a horizontal time axis along which configuration windows and update windows are shown. This is information that can be generated by the compiler after processing the intermediate representation of the program as described above.

[0046] The configuration window represents the time range during which a particular instruction is executed by a processor having a particular configuration. The update window represents the time range during which the configuration is not important or is ignored for a particular instruction.

[0047] Accordingly, the RegFile0Read instruction 452 is shown as having a first update window 412 and a configuration window 415 such that the instruction 452 is executed once per iteration cycle using the configuration named C1. During the configuration window 415, the processor executes the instruction 452 once per iteration cycle using the configuration C1.

[0048] Similarly, the RegFile1Read instruction 454 is shown as having a second update window 422 and a configuration window 425 in which the instruction 454 is executed using a configuration named C2. During the configuration window 415, the processor executes the instruction 452 once per iteration cycle using the configuration C2. In particular, the second update window 422 is much longer than the first update window 412, which means there is more flexibility for issuing configuration update instructions for the RegFile1Read instruction 454.

[0049] The multiply instruction 456 is shown as having a first configuration window 432 and a second configuration window 434. Between each of these configuration windows, the processor can execute multiple cycles of the multiplication operation, as opposed to repeating the operation as in the case of the register read instructions described above. The multiply instruction has an update window 435. For example, the update window 435 can change the data path to change the operands of the multiplication operation.

[0050] The add instruction 458 is also shown as having two configuration windows 442 and 444 during which multiple cycles of the addition operation are executed, and two update windows 443 and 445 during which one or more operands of the add instruction can be changed.

[0051] The compiler can use the update window to determine a schedule of configuration update instructions that satisfy the processor's constraints. In this example, it is assumed that the constraint is that the processor can execute configuration update instructions only at cycle X, cycle Y, and cycle Z. Additionally, the processor has an additional constraint that only two configuration update instructions can be executed in a single cycle.

[0052] Cycle X overlaps with three update windows. Cycle Y overlaps with one update window. Cycle Z also overlaps with three update windows. Therefore, the compiler can select at most two out of the three overlapping update windows on Cycle X and Cycle Z.

[0053] The compiler can choose any suitable constraint-solving procedure to determine how to schedule the five configuration updates. For example, the compiler can start with the most constrained update window and work towards the least constrained update window. If no solution can be found, the compiler can issue an error that the program cannot be executed on the processor.

[0054] As shown in Figure 4, the most constrained update window is the one that overlaps with only one of Cycle X, Cycle Y, or Cycle Z. There are four such windows that overlap with only a single cycle, namely UW412, UW435, UW443, and UW445.

[0055] Therefore, the compiler can schedule the corresponding configuration update instructions among those update windows. In other words, the compiler can schedule CU0 401 and CU3 404 during Cycle X, and can schedule CU2 403 and CU4 404 during Cycle Z.

[0056] The compiler can then remove any cycle that already has two configuration updates from consideration for the remaining update windows. Therefore, the compiler can remove Cycle X and Cycle Z from further consideration, because both of those cycles are fully occupied by the configuration update instructions.

[0057] The only remaining unassigned update window is UW422, which still overlaps with cycle Y. Therefore, the compiler can schedule CU1 402 during cycle Y, and since the scheduling of all configuration update instructions was successful, the process can end.

[0058] However, consider the case where the second update window 422 is equal to the first update window 412. In that case, since at least three configuration update instructions are scheduled for cycle X, there is no solution to the scheduling problem. In that instance, the compiler can issue an error that it was unable to successfully schedule the configuration updates.

[0059] FIG. 5 is a schematic diagram showing an example of an application specific integrated circuit (ASIC) 500, specifically a dedicated logic circuit. ASIC 500 includes a plurality of synchronous processors, referred to for simplicity as tiles. For example, ASIC 500 includes tile 502, one or more of which includes dedicated circuitry configured to perform synchronous calculations such as multiplication and addition operations. Specifically, each tile 502 may include an array of cells for computing, where each cell therein is configured to perform a mathematical operation (see, e.g., exemplary tile 200 shown in FIG. 6 and described herein). In some implementations, the tiles 502 are arranged in a grid pattern and the tiles 502 are arranged along a first dimension 501 (e.g., rows) and a second dimension 503 (e.g., columns). For example, in the example shown in FIG. 5, the tile 502 is divided into four different sections (510a, 510b, 510c, 510d), each section including 288 tiles arranged in a grid of 18 tiles high by 16 tiles wide. In some implementations, ASIC 500 shown in FIG. 5 may be understood to include a single systolic array of cells that is further divided / arranged into separate tiles, where each tile includes a subset / subarray of cells, local memory, and bus lines (see, e.g., FIG. 6).

[0060] The ASIC 500 also includes a vector processing unit 504. The vector processing unit 504 includes circuitry configured to receive an output from the tile 502 and calculate a vector calculation output value based on the output received from the tile 502. For example, in some implementations, the vector processing unit 504 includes circuitry (e.g., a multiplication circuit, an addition circuit, a shifter, and / or a memory) configured to perform an accumulation operation on the output received from the tile 502. Alternatively or additionally, the vector processing unit 504 includes circuitry configured to apply a non-linear function to the output of the tile 502. Alternatively or additionally, the vector processing unit 504 generates a normalized value, a pooled value, or both. The vector calculation output of the vector processing unit can be stored in one or more tiles. For example, the vector calculation output can be stored in a memory uniquely associated with the tile 502. Alternatively or additionally, the vector calculation output of the vector processing unit 504 can be transferred to circuitry external to the ASIC 500, e.g., as an output of a calculation. In some implementations, the vector processing unit 504 is partitioned such that each partition includes circuitry configured to receive an output from a corresponding set of tiles 502 and calculate a vector calculation output based on the output received. For example, in the example shown in FIG. 5, the vector processing unit 504 includes two rows spanning a first dimension 501, each row including 32 partitions 506 arranged in 32 columns. Each partition 506 includes circuitry (e.g., a multiplication circuit, an addition circuit, a shifter, and / or a memory) configured to perform a vector calculation as described herein based on an output (e.g., an accumulated sum) from a corresponding column of the tile 502. The vector processing unit 504 can be positioned at the center of the grid of tiles 502 as shown in FIG. 5. Other positional configurations of the vector processing unit 504 are possible.

[0061] The ASIC 500 also includes communication interfaces 508 (e.g., interfaces 508a, 508b). The communication interfaces 508 include one or more sets of serializer / deserializer (SerDes) interfaces and general-purpose input / output (GPIO) interfaces. The SerDes interface is configured to receive instructions (e.g., instructions to operate controllable bus lines as described below) and / or input data for the ASIC 500 and output data from the ASIC 500 to an external circuit. For example, the SerDes interface can be configured to transmit instructions and / or input data at a rate of 32 Gbps, a rate of 56 Gbps, or any suitable data rate via a set of SerDes interfaces included within the communication interface 508. The GPIO interface is configured to provide an interface for debugging and / or bootstrap. For example, the ASIC 500 can execute a boot program when turned on. If the program does not function, an administrator can use the GPIO interface to debug the cause of the failure.

[0062] The ASIC 500 further includes a plurality of controllable bus lines (see, e.g., FIG. 6) configured to carry data between the communication interface 508, the vector processing unit 504, and the plurality of tiles 502. The controllable bus lines include, for example, wires extending along both the first dimension 501 (e.g., rows) and the second dimension 503 (e.g., columns) of the grid. A first subset of the controllable bus lines extending along the first dimension 501 may be configured to transfer data in a first direction (e.g., to the right in FIG. 5). A second subset of the controllable bus lines extending along the first dimension 501 may be configured to transfer data in a second direction (e.g., to the left in FIG. 5). A first subset of the controllable bus lines extending along the second dimension 503 may be configured to transfer data in a third direction (e.g., upward in FIG. 5). A second subset of the controllable bus lines extending along the second dimension 503 may be configured to transfer data in a fourth direction (e.g., downward in FIG. 5).

[0063] Each controllable bus line includes a plurality of transfer elements, such as flip-flops, used to carry data along the line according to a clock signal. Transferring data via a controllable bus line may include moving data from a first transfer element of the controllable bus line to a second adjacent transfer element of the controllable bus line in each clock cycle. In some implementations, data is transferred via the controllable bus line at the rising or falling edge of the clock cycle. For example, in the first clock cycle, data present in the first transfer element (e.g., flip-flop) of the controllable bus line may be transferred to the second transfer element (e.g., flip-flop) of the controllable bus line in the second clock cycle. In some implementations, the transfer elements may be periodically spaced apart by a fixed distance from each other. For example, in some cases, each controllable bus line includes a plurality of transfer elements, and each transfer element is located within or proximate to a corresponding tile 502.

[0064] Each controllable bus line also includes a plurality of multiplexers and / or demultiplexers. The multiplexer / demultiplexer of the controllable bus line is configured to transfer data between the bus line and the components of the ASCI chip 500. For example, the multiplexer / demultiplexer of the controllable bus line can be configured to transfer data to and / or from tile 502, to and / or from vector processing unit 504, or to and / or from communication interface 508. Transferring data between tile 502, vector processing unit 504, and the communication interface may include sending a control signal to the multiplexer based on the desired data transfer to occur. The control signal can be stored in a register directly coupled to the multiplexer and / or demultiplexer. And the value of the control signal can determine, for example, which data is transferred from a source (e.g., memory within tile 502 or vector processing unit 504) to the controllable bus line, or alternatively, which data is transferred from the controllable bus line to a sink (e.g., memory within tile 502 or vector processing unit 504).

[0065] Since the controllable bus line is configured to be controlled at a local level, each tile, vector processing unit, and / or communication interface includes a unique set of control elements for operating the controllable bus line passing through that tile, vector processing unit, and / or communication interface. For example, each tile, 1D vector processing unit, and communication interface can include a corresponding set of carrier elements, multiplexers, and / or demultiplexers for controlling data transfer between that tile, 1D vector processing unit, and communication interface.

[0066] To minimize the latency associated with the operation of the ASIC chip 500, the tile 502 and the vector processing unit 504 can be arranged to reduce the distance that data travels between various components. In certain implementations, both the tile 502 and the communication interface 508 may be separated into multiple sections, and both the tile section and the communication interface section are arranged such that the maximum distance that data travels between the tile and the communication interface is reduced. For example, in some implementations, a first group of tiles 502 may be arranged in a first section on a first side of the communication interface 508, and a second group of tiles 502 may be arranged in a second section on a second side of the communication interface. As a result, the distance to the tile furthest from the communication interface can be halved compared to a configuration where all of the tiles 502 are arranged in a single section on one side of the communication interface.

[0067] Alternatively, the tiles may be arranged in a different number of sections, such as four sections. For example, in the example shown in FIG. 5, a plurality of tiles 502 of the ASIC 500 are arranged in a plurality of sections 510 (510a, 510b, 510c, 510d). Each section 510 includes a similar number of tiles 502 arranged in a grid pattern (e.g., each section 510 may include 256 tiles arranged in 16 rows and 16 columns). The communication interface 508 is divided into a plurality of sections, namely a first communication interface 508a and a second communication interface 508b, arranged on both sides of the section 510 of the tiles 502. The first communication interface 508a may be coupled to the two tile sections 510a, 510c on the left side of the ASIC chip 500 through controllable bus lines. The second communication interface 508b may be coupled to the two tile sections 510b, 510d on the right side of the ASIC chip 500 through controllable bus lines. As a result, the maximum distance that data moves to and / or from the communication interface 508 (and thus the latency associated with data propagation) can be halved compared to a mode where only a single communication interface is available. Other coupling modes of the tiles 502 and the communication interface 508 are also possible to reduce data latency. The coupling mode of the tiles 502 and the communication interface 508 can be programmed by providing control signals to the carrier elements and the multiplexers of the controllable bus lines.

[0068] In some implementations, one or more tiles 502 are configured to initiate read and write operations with respect to controllable bus lines and / or other tiles (referred to herein as "control tiles") within the ASIC 500. The remaining tiles within the ASIC 500 can be configured to perform calculations based on input data (e.g., calculate layer inferences). In some implementations, the control tiles include the same components and configurations as other tiles within the ASIC 500. The control tiles can be added as one or more surplus tiles, one or more surplus rows, or one or more surplus columns of the ASIC 500. For example, in a symmetric grid of tiles 502 where each tile 502 is configured to perform calculations on input data, one or more additional rows of control tiles can be included to handle read and write operations for the tiles 502 that perform calculations on the input data. For example, each section 510 can include 18 rows of tiles, and the last 2 rows of tiles can include control tiles. In some implementations, providing separate control tiles increases the amount of memory available in the other tiles used to perform calculations. However, separate tiles dedicated to providing control as described herein are not essential, and in some cases, separate control tiles are not provided. Instead, each tile can store in local memory instructions for initiating read and write operations for that tile.

[0069] Furthermore, each section 510 shown in FIG. 5 includes tiles arranged in 18 rows by 16 columns, although the number of tiles 502 and the arrangement of tiles within a section can vary. For example, in some cases, section 510 can include an equal number of rows and columns.

[0070] Furthermore, although FIG. 5 shows it as being divided into four sections, tile 502 may be divided into other different groups. For example, in some implementations, tile 502 may be grouped into two different sections, such as a first section above the vector processing unit 504 (e.g., closer to the top of the page shown in FIG. 5) and a second section below the vector processing unit 504 (e.g., closer to the bottom of the page shown in FIG. 5). In such an arrangement, each section may include, for example, 576 tiles arranged in a grid of 32 tiles across (along direction 501) by 18 tiles up and down (along direction 503). The sections may include other total numbers of tiles and may be arranged in arrays of different sizes. In some cases, the division between sections is characterized by the hardware features of ASIC 500. For example, as shown in FIG. 5, sections 510a, 510b may be separated from sections 510c, 510d by vector processing unit 504.

[0071] Latency can also be reduced by positioning the vector processing unit 504 relatively centrally with respect to tile section 510. In some implementations, the first half of tile 502 is arranged on the first side of the vector processing unit 504, and the second half of tile 502 is arranged on the second side of the vector processing unit 504.

[0072] For example, in the ASIC chip 500 shown in FIG. 5, the vector processing unit 504 includes two sections (e.g., two rows), each of which includes a number of segments 506 that matches the number of columns of the tile 502. Each segment 506 can be arranged and configured to receive an output such as an accumulated total from the corresponding column of the tile 502 within the section 510 of the tile. In the example shown in FIG. 5, the tile sections 510a, 510b arranged on the first side of the vector processing unit 504 (e.g., above the vector processing unit 504) can be coupled to the upper rows of the segments 506 through controllable bus lines. The tile sections 510c, 510d arranged on the second side of the vector processing unit 504 (e.g., below the vector processing unit 504) can be coupled to the lower rows of the segments 506 through controllable bus lines. Further, each tile 502 within the first half above the processing unit 504 can be arranged at the same distance from the vector processing unit 504 as each respective tile 502 within the second half below the processing unit 504, so there is no difference in the overall latency between these two halves. For example, the tile 502 in row i of the first section 510a (where the variable i corresponds to the position of the row) can be arranged at the same distance from the vector processing unit 504 as the tile 502 in row m - 1 - i of the second section of the tile (e.g., section 510c) (where m represents the total number of rows in each section, assuming the rows are incremented along the same direction in both sections).

[0073] Configuring the tile section 510 in this manner can halve the distance that data moves to and / or from the vector processing unit 504, compared to an arrangement where the vector processing unit 504 is placed at the edge (e.g., bottom) of all tiles 502. For example, the latency associated with receiving the accumulated sum of a column of tiles 502 from section 510a can be halved compared to the latency associated with receiving the accumulated sum of a column of tiles 502 from sections 510a and 510c. The coupling pattern of the tiles 502 and the vector processing unit 504 can be programmed by providing control signals to a multiplexer of carrier elements and controllable bus lines.

[0074] During operation of the ASIC chip 500, the activation input can be shifted between tiles. For example, the activation input can be shifted along the first dimension 501. Additionally, the output from the calculations performed by the tiles 502 (e.g., the output of the calculations performed by the calculation array within the tile 502) can be shifted along the second dimension 503 between tiles.

[0075] In some implementations, the controllable bus lines can be physically hardwired to skip tiles 502 in the data in order to reduce the latency associated with the operation of the ASIC chip 500. For example, the output of a calculation performed by a first tile 502 can be shifted along a second dimension 503 of the grid to a second tile 502 that is located at least one tile away from the first tile 502, skipping the tiles in between. In another example, an activation input from a first tile 502 can be shifted along a first dimension 501 of the grid to a second tile 502 that is located at least one tile away from the first tile 502, skipping the tiles in between. By skipping at least one tile when shifting the activation input or output data, the overall data path length can be reduced so that the data is transferred faster (e.g., there is no need to use clock cycles to store the data in the skipped tiles), and the latency is reduced.

[0076] In an exemplary implementation, each tile 502 within each column of section 510a can be configured to pass output data to vector processing unit 504 along a second dimension 503 through controllable bus lines. Tiles 502 within each column can further be configured to pass data to vector processing unit 504 by skipping the next adjacent tile (e.g., through physical hardwiring of controllable bus lines between tiles). That is, the tile 502 at position (i,j)=(0,0) in the first section 510a (where variable i corresponds to the row position and variable j corresponds to the column position) may be hardwired to pass output data to the tile 502 at position (i,j)=(2,0), and similarly, the tile 502 at position (i,j)=(2,0) in the first section 510a may be hardwired to pass output data to the tile 502 at position (i,j)=(4,0), and so on. The last tile not skipped (e.g., the tile 502 located at position (i,j)=(16,0)) passes the output data to vector processing unit 504. For a section 510 having 18 rows of tiles, such as the example shown in FIG. 5, tile skipping ensures that all tiles within section 510 are at most 9 "tile hops" away from vector processing unit 504, thus improving the performance of ASIC chip 500 by reducing the data path length and the resulting data latency by half.

[0077] In another exemplary implementation, each tile 502 within each row of sections 510a, 510c and within each row of sections 510b, 510d can be configured to pass activation inputs along the first dimension 501 through controllable bus lines. For example, some tiles within sections 510a, 510b, 510c, 510d can be configured to pass activation inputs towards the center of the grid 500 or towards the communication interface 508. The tiles 502 within each row can further be configured to skip adjacent tiles, for example, by hardwiring controllable bus lines between the tiles. For example, a tile 502 at position (i,j)=(0,0) in the first section 510a (where variable i corresponds to the row position and variable j corresponds to the column position) may be configured to pass an activation input to the tile 502 at position (i,j)=(0,2), and similarly, a tile 502 at position (i,j)=(0,2) in the first section 510a may be configured to pass an activation input to the tile 502 at position (i,j)=(0,4), and so on. In some cases, the last tile that is not skipped (e.g., the tile 502 located at position (i,j)=(0,14)) does not pass the activation input to another tile.

[0078] Similarly, the skipped tiles can pass the activation input in the opposite direction. For example, tile 502 at position (i,j)=(0,15) in the first section 510a (where variable i corresponds to the row position and variable j corresponds to the column position) may be configured to pass the activation input to tile 502 at position (i,j)=(0,13). Similarly, tile 502 at position (i,j)=(0,13) in the first section 510a may be configured to pass the activation input to tile 502 at position (i,j)=(0,11), and so on. In some cases, the last tile that is not skipped (for example, tile 502 located at position (i,j)=(0,1)) does not pass the activation input to another tile. By skipping tiles, in some implementations, it is possible to improve the performance of the ASIC chip 500 by reducing the data path length and the resulting data latency by half.

[0079] As described herein, in some implementations, one or more of the tiles 502 are dedicated to storing control information. That is, a tile 502 dedicated to storing control information is not involved in performing calculations on input data such as weight inputs and activation inputs. The control information may include, for example, control data for configuring controllable bus lines during operation of the ASIC chip 500 so that data can be moved around the ASIC chip 500. The control data may be provided to the controllable bus lines in the form of control signals for controlling the carrier elements and multiplexers of the controllable bus lines. The control data defines whether a particular carrier element of the controllable bus line passes data to the next carrier element of the controllable bus line so that data is transferred between tiles according to a predetermined schedule. The control data additionally defines whether data is transferred from or to the bus line. For example, the control data may include control signals that direct a multiplexer to transfer data from the bus line to memory and / or other circuitry within the tile. In another example, the control data may include control signals that direct a multiplexer to transfer data from memory and / or circuitry within the tile to the bus line. In another example, the control data may include control signals that direct a multiplexer to transfer data between the bus line and the communication interface 508 and / or between the bus line and the vector processing unit 504. Alternatively, as disclosed herein, dedicated control tiles are not used. Rather, in such cases, the local memory of each tile stores control information for that particular tile.

[0080] FIG. 6 shows an example of a tile 600 for use in an ASIC chip 500. Each tile 600 includes a local memory 602 and a computing array 604 coupled to the memory 602. The local memory 602 includes physical memory located near the computing array 604. The computing array 604 includes a plurality of cells 606. Each cell 606 of the computing array 604 includes circuitry configured to perform a computation (e.g., a multiply-accumulate operation) based on data inputs such as an activation input and a weight input to the cell 606. Each cell can perform a computation (e.g., a multiply-accumulate operation) at the period of a clock signal. The computing array 604 can have more rows than columns, more columns than rows, or an equal number of rows and columns. For example, in the example shown in FIG. 6, the computing array 604 includes 64 cells arranged in 8 rows and 8 columns. In particular, other computing array sizes are possible, such as a computing array having 16 cells, 32 cells, 128 cells, or 256 cells. Each tile can include the same number of cells and / or a computing array of the same size. Then, the total number of operations that can be executed in parallel for the ASIC chip depends on the total number of tiles having computing arrays of the same size within the chip. For example, for the ASIC chip 500 shown in FIG. 5 that includes approximately 1150 tiles, this means that approximately 72,000 computations can be executed in parallel per cycle. Examples of clock speeds that can be used include, but are not limited to, 225 MHz, 500 MHz, 750 MHz, 1 GHz, 1.25 GHz, 1.5 GHz, 1.75 GHz, or 2 GHz. The computing array 604 of each individual tile is a subset of a larger systolic array of tiles, as shown in FIG. 5.

[0081] The memory 602 included in the tile 600 may include a random access memory (RAM), such as SRAM. Alternatively, other memories may be used. Each memory 602 may be configured to store one nth of the entire memory associated with the n tiles 502 of the ASIC chip shown in FIG. 5. The memory 602 may be provided as a single chip or in multiple chips. For example, the memory 602 shown in FIG. 6 is provided as four single-port SRAMs, each of which is coupled to the compute array 604. Alternatively, the memory 602 may be provided as, among other configurations, two single-port SRAMs or eight single-port SRAMs. The total capacity of the memory may be, but is not limited to, for example, 16 kB, 32 kB, 64 kB, or 128 kB after error correction coding. By providing the physical memory 602 locally to the compute array, the wiring density for the ASIC 500 can be significantly reduced in some implementations. In contrast to being provided locally as described herein, in an alternative configuration where the memory is centralized within the ASIC 500, wires may be required for each bit of the memory bandwidth. The total number of wires required to cover each tile of the ASIC 500 greatly exceeds the available space within the ASIC 100. In contrast, using dedicated memory provided for each tile, the total number required to cover the area of the ASIC 500 can be significantly reduced.

[0082] Tile 600 also includes controllable bus lines. The controllable bus lines can be classified into multiple different groups. For example, the controllable bus lines can include a first group of general-purpose controllable bus lines 610 configured to transfer data between tiles in each of the major directions. That is, the first group of controllable bus lines 610 includes a bus line 610a configured to transfer data in a first direction (referred to as "east" in FIG. 6) along the first dimension 501 of the tile grid, a bus line 610b configured to transfer data in a second direction (referred to as "west" in FIG. 6) along the first dimension 101 of the tile grid, where the second direction is opposite to the first direction, a bus line 610c configured to transfer data in a third direction (referred to as "north" in FIG. 6) along the second dimension 103 of the tile grid, and a bus line 610d configured to transfer data in a fourth direction (referred to as "south" in FIG. 6) along the second dimension 103 of the tile grid, where the fourth direction is opposite to the third direction. The general-purpose bus lines 610 can be configured to carry control data, activation input data, data from and / or to the communication interface, data from and / or to the vector processing unit, and data to be stored and / or used by tile 600 (e.g., weight input). Tile 600 can include one or more control elements 621 (e.g., flip-flops and multiplexers) for controlling the controllable bus lines and thus for routing data to and / or from tile 600 and / or memory 602.

[0083] The controllable bus lines may also include a second group of controllable bus lines, herein called the compute array partial sum bus lines 620. The compute array partial sum bus lines 620 may be configured to carry data output from the computations performed by the compute array 604. For example, as shown in FIG. 6, the bus lines 620 may be configured to carry partial sum data obtained from the rows within the compute array 604. In such a case, the number of bus lines 620 will match the number of rows within the array 604. For example, for an 8×8 compute array, there are eight partial sum bus lines 620 each coupled to the output of the corresponding row within the compute array 604. The compute array output bus lines 620 may further be configured to couple to another tile within the ASIC chip, for example as an input to a compute array of another tile within the ASIC chip. For example, the array partial sum bus lines 620 of tile 600 may be configured to receive an input (e.g., partial sum 620a) to the compute array of a second tile located at least one tile away from tile 600. The output of the compute array 604 is then added to the partial sum line 620 to produce a new partial sum 620b which may be the output from tile 600. The partial sum 620b may then be passed to another tile or alternatively to the vector processing unit. For example, each bus line 620 may be coupled to a corresponding segment (such as segment 506 of FIG. 5) of the vector processing unit.

[0084] As described with respect to FIG. 5, a controllable bus line may include circuitry such as a transport element (e.g., a flip-flop) configured to enable data to be transported along the bus line. In some implementations, each controllable bus line includes a corresponding transport element for each tile. As further described with respect to FIG. 5, a controllable bus line may include circuitry such as a multiplexer configured to enable data to be transferred between different tiles, vector processing units, and the communication interfaces of the ASIC chips. The multiplexer may be placed anywhere there is a source or sink of data. For example, in some implementations, as shown in FIG. 6, a control circuit 621 such as a multiplexer is placed at the intersection of controllable bus lines (e.g., at the intersection of general-purpose bus lines 610a and 610d, at the intersection of general-purpose bus lines 610a and 610c, at the intersection of general-purpose bus lines 610b and 610d, and / or at the intersection of general-purpose bus lines 610b and 610c). The multiplexer at the intersection of the bus lines may be configured to transfer data between the bus lines at the intersection. Thus, by appropriate operation of the multiplexer, it may be possible to change the direction in which data moves through the controllable bus line. For example, data moving along the first dimension 101 on the general-purpose bus line 610a may be transferred to the general-purpose bus line 610d, so that the data instead moves along the second dimension 103. In some implementations, the multiplexer may be placed adjacent to the memory 602 of the tile 600 such that data may be transferred to and / or from the memory 602.

[0085] Embodiments of the subject matter and the functional operations described in this specification can be implemented in digital electronic circuitry, in tangibly embodied computer software or firmware, in computer hardware including the structures disclosed in this specification and their structural equivalents, or in one or more combinations of them. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., as one or more modules of computer program instructions encoded on a tangible non-transitory storage medium for execution by, or to control the operation of, a data processing apparatus. The computer storage medium can be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or a combination of one or more of them. Alternatively or in addition, the program instructions can be encoded on an artificially generated propagated signal, e.g., a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode information for transmission to an appropriate receiver apparatus for execution by a data processing apparatus.

[0086] The term “data processing apparatus” refers to data processing hardware and encompasses all kinds of devices, devices, and machines for processing data, including by way of example a programmable processor, a computer, or multiple processors or computers. The apparatus can also be, or further include, special purpose logic circuitry, e.g., an FPGA (Field Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The apparatus optionally includes, in addition to hardware, code that creates an execution environment for computer programs, e.g., processor firmware, a protocol stack, a database management system, an operating system, or a combination of one or more of them.

[0087] A computer program, also called a program, software, software application, app, module, software module, script, or code, or sometimes described as such, may be written in any form of programming language, including a compiled or interpreted language, or a declarative or procedural language, and may be deployed in any form, including as a stand-alone program or including as a module, component, subroutine, or other unit suitable for use in a computing environment. The program may or may not correspond to a file in a file system. The program may be stored in a portion of a file that holds one or more scripts stored in a markup language document, in a single file dedicated to the program, or in multiple cooperating files, such as files that hold one or more modules, subprograms, or portions of code. The computer program may be deployed to be executed on one computer or on multiple computers located at one location or distributed across multiple locations and interconnected by a data communication network.

[0088] Means that software, firmware, hardware, or a combination thereof that causes the system to perform an operation or activity is installed on the system so that one or more computer systems are configured to perform a particular operation or activity. Means that when one or more computer programs are configured to perform a particular operation or activity and are executed by a data processing apparatus, the one or more programs include instructions that cause the apparatus to perform the operation or activity.

[0089] As used herein, "engine" or "software engine" refers to an input / output system implemented in software that provides an output different from the input. An engine can be an encoded block of functionality, such as a library, platform, software development kit ("SDK"), or object. Each engine can be implemented on any suitable type of computing device, such as a server, mobile phone, tablet computer, notebook computer, music player, e - book reader, laptop or desktop computer, PDA, smartphone, or other fixed or portable device that includes one or more processors and a computer - readable medium. Additionally, two or more of the engines can be implemented on the same computing device or on different computing devices.

[0090] The processes and logical flows described herein can be implemented by one or more programmable computers that execute one or more computer programs to perform functions by acting on input data to generate output. The processes and logical flows can also be implemented by dedicated logic circuitry, such as an FPGA or ASIC, or by a combination of dedicated logic circuitry and one or more programmed computers.

[0091] A computer suitable for the execution of a computer program may be based on a general-purpose microprocessor or a special-purpose microprocessor or both, or any other kind of central processing unit. Generally, the central processing unit receives instructions and data from a read-only memory or a random access memory or both. Indispensable elements of a computer are a central processing unit for executing or running instructions, and one or more memory devices for storing instructions and data. The central processing unit and the memory may be augmented by, or incorporated in, dedicated logic circuitry. Generally, a computer includes one or more mass storages for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or is operatively coupled to receive data from, or transfer data to, or both, such devices. However, a computer need not have such devices. Moreover, a computer may be incorporated in another device, such as, by way of example, a cellular phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device, such as a universal serial bus (USB) flash drive.

[0092] Computer-readable media suitable for storing computer program instructions and data include, by way of example, all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices, e.g., EPROM, EEPROM, and flash memory devices, magnetic disks, e.g., internal hard disks or removable disks, magneto-optical disks, and CD-ROM disks and DVD-ROM disks.

[0093] To enable interaction with a user, embodiments of the subject matter described herein may be implemented on a computer having a display device, such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor, for displaying information to the user, as well as a keyboard and a pointing device, such as a mouse, trackball, or presence-sensing display or other surface by which the user can provide input to the computer. Other types of devices may also be used to enable interaction with the user. For example, feedback provided to the user may be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback. Input from the user may be received in any form, including acoustic input, speech input, or tactile input. Additionally, the computer may interact with the user by sending and receiving documents from the devices used by the user, for example, by sending a web page to a web browser in response to a request received from the web browser on the user's device, or by sending a text message or other form of message to a smartphone that runs a messaging application and receives a response message from the user as a reply.

[0094] Embodiments of the subject matter described in this specification may be implemented in a computing system that includes back-end components, such as a data server, or includes middleware components, such as an application server, or includes front-end components, such as a graphical user interface, a web browser, or a client computer having an app with which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end components, middleware components, or front-end components. The components of the system may be interconnected by any form or medium of digital data communication, such as a communication network. Examples of communication networks include local area networks (LANs) and wide area networks (WANs), such as the Internet.

[0095] Embodiment 1 is receiving an intermediate representation of a program that defines a plurality of operations to be performed by respective components of a synchronous processor, the intermediate representation assigning to each operation of the plurality of operations a respective clock cycle value by which the operation is scheduled to be executed by the synchronous processor; processing the intermediate representation to generate, for each operation in the intermediate representation that requests a hardware configuration update, a respective update window, the update window defining a time range during which a configuration update command can be executed to perform the hardware configuration update; obtaining one or more configuration constraints for the synchronous processor; generating and scheduling configuration update commands that each occur during one of the update windows according to the configuration constraints of the synchronous processor; and a method comprising:

[0096] Embodiment 2 is the method of Embodiment 1, wherein the configuration constraint comprises a maximum number of configuration update commands that the synchronous processor can execute in a single cycle.

[0097] Embodiment 3 is the method of Embodiment 2, and the configuration constraint comprises a limited number of cycles in which a configuration update command can be executed.

[0098] Embodiment 4 is any one of the methods of Embodiments 1 to 3, and the step of generating and scheduling a configuration update command from an update window and a configuration constraint comprises the step of allocating the configuration update command to the most restricted update window before allocating the configuration update command to other update windows.

[0099] Embodiment 5 is any one of the methods of Embodiments 1 to 4, and one or more of the configuration update commands change the registers of the read operation.

[0100] Embodiment 6 is any one of the methods of Embodiments 1 to 5, and one or more of the configuration update commands change the multiplexer arguments in the data path of the synchronous processor.

[0101] Embodiment 7 is any one of the methods of Embodiments 1 to 6, and the synchronous processor is configured to execute each operation using the same configuration when no configuration update command is issued.

[0102] Embodiment 8 is any one of the methods of Embodiments 1 to 7, and further comprises the step of executing an operation using a synchronous processor and the step of updating the configuration of the synchronous processor according to a configuration update command.

[0103] Embodiment 9 is the step of receiving a program that defines a plurality of operations to be executed by a plurality of respective components of a synchronous processor, and for each component of the plurality of components, the step of obtaining respective clock latency values for each respective operation of the program to be executed by the component, generating clock timing constraints for executing a program based on respective clock latency values for respective operations of the program; generating an intermediate representation of a program that assigns respective clock cycle values to respective operations of a plurality of operations based on the clock timing constraints; The method includes the following steps.

[0104] Embodiment 10 is the method of Embodiment 9, wherein the operation further includes: processing the intermediate representation to generate respective update windows for respective operations in the intermediate representation that request a hardware configuration update, the update window defining a time range during which a configuration update command can be executed to perform the hardware configuration update; obtaining one or more configuration constraints for a synchronous processor; generating and scheduling configuration update commands that each occur during one of the update windows according to the configuration constraints of the synchronous processor; The method includes the following steps.

[0105] Embodiment 11 is the method of Embodiment 10, wherein the configuration constraint includes a maximum number of configuration update commands that the synchronous processor can execute in a single cycle.

[0106] Embodiment 12 is the method of Embodiment 11, wherein the configuration constraint includes a limited number of cycles in which the configuration update command can be executed.

[0107] Embodiment 13 is the method of Embodiment 10, wherein the step of generating and scheduling configuration update commands from the update window and the configuration constraint includes allocating the configuration update commands to the most restricted update window before allocating the configuration update commands to other update windows.

[0108] Embodiment 14 is the method of any one of Embodiments 9 to 13, wherein the program is not affected by latency.

[0109] Embodiment 15 is the method of Embodiment 14, and the program does not define the order of operations or the timing between operations.

[0110] Embodiment 16 is the method of Embodiment 15, and the program does not define latency information for components of the synchronous processor.

[0111] Embodiment 17 is the method of any one of Embodiments 9 to 16, and the intermediate representation of the program can assign negative clock cycle values to each of a plurality of operations.

[0112] Embodiment 18 is the method of any one of Embodiments 9 to 17, and the operation further includes a step of generating dependency relationship information between operations from variables in the input program.

[0113] Embodiment 19 is a system including one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, are operable to cause the one or more computers to execute the method of any one of the preceding embodiments.

[0114] Embodiment 20 is a computer storage medium encoded with a computer program, and the program includes instructions that, when executed by a data processing device, are operable to cause the data processing device to execute the method of any one of Embodiments 1 to 18.

[0115] This specification includes many details of specific implementations, which should not be regarded as restrictions on the scope of the invention or what can be claimed, but rather as descriptions of features that may be specific to particular embodiments of a particular invention. Some features described herein in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented separately in multiple embodiments, or in any suitable sub-combination. Moreover, features are described above as acting in certain combinations and may thus be initially claimed as such, but one or more features from the claimed combination may in some cases be deleted from the combination, and the claimed combination may be directed to a sub-combination or a variation of a sub-combination.

[0116] Similarly, operations are shown in the drawings in a particular order, which should not be understood as requiring that such operations be performed in the particular order or sequential order shown, or that all of the shown operations be performed, in order to achieve a desired result. In some situations, multitasking and parallel processing may be advantageous. Moreover, the separation of various system modules and components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be understood that the program components and systems described may generally be integrated together in a single software product, or packaged into multiple software products.

[0117] Specific embodiments of the subject matter have been described. Other embodiments are within the scope of the following claims. For example, the acts recited in the claims may be performed in a different order and still achieve a desired result. As an example, the processes shown in the accompanying drawings do not necessarily require the particular order or sequential order shown in order to achieve a desired result. In some cases, multitasking and parallel processing may be advantageous.

Explanation of Symbols

[0118] 412 Update Window 415 Configuration Window 422 Update Window 425 Configuration Window 432 First Configuration Window 434 Second Configuration Window 435 Update Window 442 Configuration Window 443 Update Window 444 Configuration Window 445 Update Window 452 RegFile0Read Command 454 RegFile1Read Command 456 Multiplication Command 458 Addition Command 501 First Dimension 502 Tile 503 Second Dimension 504 Vector Processing Unit 506 Division 508 Communication Interface 510 Section 600 Tile 602 Local Memory 604 Computing Array 606 Cell 610 Bus Line 620 Bus Line 621 Control Element

Claims

Claim 1 Receiving an intermediate representation of a program that defines a plurality of operations to be performed by respective components of a synchronous processor, wherein the intermediate representation assigns a respective clock cycle value to each of the plurality of operations such that the operation is scheduled to be executed by the synchronous processor; Processing the intermediate representation to generate respective update windows for each operation in the intermediate representation that requests a hardware configuration update, wherein the update window defines a time range during which a configuration update instruction can be executed to perform the hardware configuration update; Obtaining one or more configuration constraints for the synchronous processor; Generating and scheduling configuration update instructions that each occur during one of the update windows according to the configuration constraints of the synchronous processor. A method comprising the above steps.

Citation Information

Patent Citations

  • Management device, method for controlling management device, and program

    JP2016184274A

  • Method and apparatus for computing

    US20060004997A1

  • Reconfigurable instruction cell array

    US20100122105A1

  • Systems and methods for instruction entity allocation and scheduling on multi-processors

    US20140122848A1

  • Compilation for synchronous processor

    WO2021035187A1