Rotating data blocks

CN120390922APending Publication Date: 2025-07-29GRAPHCORE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202380087658.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-20
Filing Date
2023-11-20
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to meet the timing requirements of relatively short clock cycles when performing byte-byte rotation operations of data blocks in high-performance processing devices, and the wiring of large multiplexer arrays is complex, resulting in high area cost and increased power consumption.

Method used

The logic circuit design is adopted, and the data blocks are rotated layer by layer by layer through input data arrays, multi-layer multiplexer arrays and control signal generators, which avoids the use of large multiplexer arrays, and uses smaller multiplexers and pipeline registers to meet timing requirements.

Benefits of technology

It realizes efficient and low-cost data block rotation operation in high-performance processing equipment, reducing wiring complexity, and reducing area and power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390922A_ABST
    Figure CN120390922A_ABST
Patent Text Reader

Abstract

The execution unit performs a byte-by-byte rotation operation of an input data block. The input data array receives an input data block. The two first layer multiplexer arrays each receive a first layer data block including a respective subset of bytes of an input data block and a first layer control signal, and rotate the first layer data block by an amount indicated by the first layer control signal. The second layer multiplexer array receives a second control signal and selects between respective bytes of the first rotated first layer data block and the second rotated first layer data block based on the second control signal. The operation unit further comprises a control signal generator used for generating a first layer control signal and a second layer control signal according to the received computer program instructions. Thus, the result of the smaller block rotation is used as part of the result of the larger block rotation, avoiding a large multiplexer array with complex wiring.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to a logic circuit for performing byte-by-byte rotation of data blocks. Background Art

[0002] A processing device may include a running unit and a memory. The running unit is capable of running one or more program threads to perform operations on data loaded from the memory to generate results, and then storing the results in the memory. The results may be subject to subsequent processing by the running unit, or may be dispatched from the processing device. Summary of the Invention

[0003] Data blocks loaded from the memory by the processing device may have one of a plurality of predetermined block sizes, each block size having a fixed number of bytes, such as 4, 8, or 16 bytes. In various contexts, the running unit may need to rotate data blocks of various block sizes stored in the memory by any number (up to the block size) of bytes. As the name implies, "rotating" a data block by any number of bytes includes moving bytes from the front (i.e., having the highest index) of the block to the back. For example, rotating a 4-byte data block [A, B, C, D] of 1 block is [B, C, D, A]. For example, block rotation may be required during the process of aligning misaligned data.

[0004] Such rotation operations need to meet relevant timing requirements, which may be problematic in the context of high-performance processing devices with relatively short clock cycles (e.g., processors for machine learning applications).

[0005] The object of the present disclosure is to provide a logic circuit that can perform rotation operations within the timing requirements of such high-performance processing devices. Another object of the present disclosure is to provide a means for performing rotation operations in a manner that reduces area cost and minimizes power consumption.

[0006] According to a first aspect of the present disclosure, there is provided a running unit configured to run computer program instructions to perform a byte-by-byte rotation operation on an input data block. The running unit includes:

[0007] A logic circuit, including:

[0008] An input data array for receiving an input data block including N bytes;

[0009] Two first-layer multiplexer arrays, each first-layer multiplexer array being configured to:

[0010] Receive a first-layer data block, the first-layer data block including a corresponding byte subset of the input data block;

[0011] Receive first-layer control signals;

[0012] Rotate the first - layer data block by an amount indicated by the first - layer control signal;

[0013] Two first - layer multiplexer arrays, configured to output a first rotated first - layer data block and a second rotated first - layer data block respectively;

[0014] A second - layer multiplexer array, configured to receive a second control signal, the second - layer multiplexer array including N multiplexers, each multiplexer being configured to select between corresponding bytes of the first rotated first - layer data block and the second rotated first - layer data block based on the second control signal to output a rotated second - layer data block, and

[0015] A control signal generator, configured to generate the first - layer control signal and the second - layer control signal based on the received computer program instructions.

[0016] Each first - layer data block may include N / 2 bytes. N may be equal to 8.

[0017] The logic circuit may be extended by adding additional layers to perform rotation of larger input data blocks. In one example, the input data array is configured to receive an input data block including M bytes, where M > N, suitably where M = 2N. M may be equal to 16. The logic circuit may include four first - layer multiplexer arrays to output first through fourth rotated first - layer blocks. The logic circuit may include two second - layer multiplexer arrays, the first of the second - layer multiplexer arrays being configured to receive the first rotated first - layer data block and the second rotated first - layer block and output a first rotated second - layer data block, the second of the second - layer multiplexer arrays being configured to receive the third rotated first - layer block and the fourth rotated first - layer block and output a second rotated second - layer data block. The logic circuit may include a third - layer multiplexer array configured to receive a third - layer control signal. The third - layer multiplexer array may include M multiplexers, configured to select between corresponding bytes of the first rotated second - layer data block and the second rotated second - layer data block based on the third control signal to output a rotated third - layer data block. The control signal generator may be configured to generate the third - layer control signal.

[0018] The blocks herein may include a plurality of bytes that are a power of 2. For example, N and / or M may be a power of 2.

[0019] The logic circuit may include an intermediate result array, the intermediate result array being configured to receive the output of the first - layer multiplexer array. The second - layer multiplexer array may receive the first rotated first - layer data block and the second rotated first - layer data block from the intermediate result array. In an example where the logic circuit includes additional layers, additional intermediate result arrays may be provided between consecutive layers.

[0020] Each first - layer multiplexer array may include a plurality of S:1 multiplexers, suitably S S:1 multiplexers. S may be equal to the size of the byte subset of the input block. The input j of the i - th S:1 multiplexer of the array may be connected to the byte (i + j) mod S of the subset of the input data array corresponding to the respective byte subset.

[0021] The second - layer control signal may include a bit - mask of N bits, and each bit in the bit - mask serves as a control signal for a respective one of the N multiplexers in the second multiplexer array. The logic circuit may include circuitry for splitting the bit - mask and providing the respective bits to the respective multiplexers.

[0022] The control - signal generator may be configured to rotate the bit - mask, where the amount of rotation of the bit - mask results in a rotated second - layer data block that is rotated by the same amount at the output. The bit - mask may be 0xF0.

[0023] The control - signal generator may be configured to rotate the bit - mask by selecting a stored rotated bit - mask from a look - up table. The control - signal generator may include circuitry, such as a shift - bit circuit, configured to rotate the bit - mask.

[0024] The third - layer control signal may be a bit - mask of M bits, where the third - layer multiplexer array is controlled similarly to the second - layer control signal. The bit - mask of the third - layer control signal may be 0xFF00.

[0025] The control - signal generator may generate a second - layer control signal that causes the second - layer multiplexer array to function as a pass - through. Thus, the second - layer multiplexer array may output a first rotated first - layer data block and a second rotated first - layer data block. In an example including additional layers, the control - signal generator may generate control signals to cause the additional layers to function as pass - throughs. For example, the control - signal generator may generate a third - layer control signal that causes the third - layer to function as a pass - through.

[0026] The logic circuit may include one or more clock gating configured to disable one or more elements of the logic circuit. The logic circuit may include a plurality of data - path channels, suitably N data - path channels. Each clock gating may be configured to disable one or more data - path channels. Each clock gating may be configured to disable N / 2 or M / 4 data - path channels, suitably corresponding to an input of one in the first - layer multiplexer array. The clock gating may be configured to disable one of the first - layer multiplexer array and the second - layer multiplexer array. The control - signal generator may be configured to generate clock - gating control signals to control one or more clock gating.

[0027] The logic circuit may include pipeline registers disposed between the first - layer multiplexer array and the second - layer multiplexer array.

[0028] The computer program instructions can be rotation instructions, configured to rotate an input data block. The computer program instructions can include a plurality of operations, where the rotation operation is one of the plurality of operations. The computer program instructions can include a packing instruction, configured to copy a sequence of consecutive bytes from a first location in a first data block to a second location in a second data block. The rotation operation can align the bytes of the first data block to an output location. The computer program instructions can be extraction instructions, configured to extract a sequence of consecutive bytes from a concatenation of the first data block and the second data block.

[0029] The execution unit can generate a first layer of control signals, a second layer of control signals, and optionally a third layer of control signals based on values indicated by the computer program instructions. The values can be indicated in one or more operands of the computer program instructions, in the opcode of the computer program instructions, or can be read from one or more registers associated with the computer program instructions. These values can include a rotation amount and / or a block size. These values can include one or more values, and the rotation amount and / or the block size can be calculated by the execution unit based on these values.

[0030] According to a second aspect of the present disclosure, there is provided a processing unit including the execution unit defined in the first aspect. The processing unit can be a tile processor. The processing unit can include a local memory. The execution unit can receive an input data block from the local memory.

[0031] According to a third aspect of the present disclosure, there is provided a processing device including the processing unit defined in the second aspect. The processing device can include a plurality of processing units. At least one of the processing units can include an execution unit. The processing units can communicate via a switching fabric that implements time-deterministic switching.

[0032] According to a fourth aspect of the present disclosure, there is provided a method implemented in an execution unit, the method including:

[0033] Receiving an input data block including N bytes;

[0034] Providing a first layer of data blocks including corresponding byte subsets of the input data block to a first first-layer multiplexer array and a second first-layer multiplexer array;

[0035] Providing first layer control signals to the first first-layer multiplexer array and the second first-layer multiplexer array;

[0036] Rotating the first layer of data blocks by an amount indicated by the control signals to output a first rotated first layer of data blocks and a second rotated first layer of data blocks;

[0037] Provide the first rotated first - layer data block and the second rotated first - layer data block to a second - layer multiplexer array, the second - layer multiplexer array including N multiplexers,

[0038] Provide a second control signal to the second - layer multiplexer array to select between corresponding bytes of the first rotated first - layer data block and the second rotated first - layer data block based on the second control signal.

[0039] Further optional features of the method of the fourth aspect are defined above with respect to the first, second, and third aspects and may be combined in any combination. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] For a better understanding of the present invention and to show how it may be implemented, reference will now be made, by way of example only, to the accompanying drawings, in which:

[0041] Figure 1A is a schematic block diagram of a processor in which an example of the present disclosure is implemented;

[0042] Figure 1B is a schematic diagram of an example of a processor chip in which an example of the present disclosure is implemented;

[0043] Figure 2 is a schematic block diagram of a simple example logic circuit;

[0044] Figure 3 is a schematic block diagram of an example logic circuit of the present disclosure;

[0045] Figure 4 is shown in more detail Figure 3 a schematic block diagram of a first processing stage of an example logic circuit;

[0046] Figure 5 is shown in more detail Figure 3 a schematic block diagram of a second processing stage of an example logic circuit;

[0047] Figure 6 is shown in more detail Figure 3 a schematic block diagram of a third processing stage of an example logic circuit;

[0048] Figure 7 is another schematic block diagram of an example logic circuit;

[0049] Figure 8 is another schematic block diagram of an example logic circuit;

[0050] Figure 9 is an example code sample showing a packing instruction;

[0051] Figure 10 is an example code sample showing an extraction instruction.

[0052] In the drawings, corresponding reference numerals denote corresponding components. Those skilled in the art will appreciate that the elements in the drawings are shown for simplicity and clarity and are not necessarily drawn to scale. For example, the dimensions of some of the elements in the drawings may be exaggerated relative to other elements to help improve understanding of the various examples. Additionally, commonly understood elements that are useful or necessary in a commercially viable example are often not depicted so as to facilitate a less obstructed view of these various examples. Detailed Description

[0053] Generally, examples of the present disclosure provide a logic circuit for performing a byte-by-byte rotation of a data block. The logic circuit is logically arranged in layers, where each layer is configured to perform operations on successively larger sized blocks. The result of the previous layer provides an input to the subsequent layer such that the result of the rotation of the smaller block is used as a partial result for computing the rotation of the larger block. In some examples, the logic circuit is incorporated into a processing unit of a processing device (such as a tile processor of a processor having multiple tiles), for example, in a execution unit of the processing unit.

[0054] Advantageously, examples of the present disclosure provide a means for rotating relatively large sized blocks, which avoids large multiplexer arrays with complex wiring. Thus, smaller and faster multiplexers are employed and the need for routing a large number of long data paths is avoided. This in turn helps to keep the timing paths short. Additionally, examples herein assist in minimizing the area cost of processor features by employing hardware resources that can be reused across rotation operations of all sizes (as opposed to requiring different hardware resources for different block sizes).

[0055] The examples are implemented in a processing unit, which may take the form of processor 4, reference Figure 1A Processor 4 is described in more detail. In some examples, processor 4 may take the form of tile 4 of a multi-tile processing device. Examples of such multi-tile processing devices are described in more detail in our earlier application U.S. Patent Application No. US 2020 / 0319861A1, which is incorporated herein by reference.

[0056] Reference Figure 1A , Figure 1AShows an example of a processor 4, including details of the execution unit 18 and the context register 26. The illustrated processor 4 includes a weight register file 26W and can thus be particularly suitable for machine learning applications, where those models are trained by adjusting the weights of the machine learning models. However, the examples of applications are not limited to machine learning applications but are more widely applicable. In addition, the described processor 4 is a multi-threaded processor capable of running M threads simultaneously. The processor 4 is capable of supporting the execution of M worker threads and one supervisor thread, where the worker threads perform arithmetic operations on data to generate results, and the supervisor thread coordinates the worker threads and controls the synchronization, sending, and receiving functions of the processor 4.

[0057] The processor 4 includes a corresponding instruction buffer 53 for each of the M threads that can run simultaneously. The context register 26 includes a corresponding main register file (MRF) 26M for each of the M worker contexts and one supervisor context. The context register also includes a corresponding auxiliary register file (ARF) 26A for at least one of the worker contexts. The context register 26 also includes a common weight register file (WRF) 26W that all currently running worker threads can access to read from the WRF. The WRF can be associated with the supervisor context because the supervisor thread is the only thread that can write to the WRF. The context register 26 can also include a corresponding set of control status registers 26CSR for each of the supervisor context and the worker contexts. The execution unit 18 includes a main execution unit 18M and an auxiliary execution unit 18A. The main execution unit 18M includes a load-store unit (LSU) 55 and an integer arithmetic logic unit (IALU) 56. The auxiliary execution unit 18A includes at least a floating-point arithmetic unit (FPU).

[0058] In each of the J interleaved time slots S0…SJ-1, the scheduler 24 controls the fetch stage to fetch at least one instruction of the corresponding thread from the instruction memory 11 into the corresponding one of the J instruction buffers 53 corresponding to the current time slot. In the example, each time slot is one run loop of the processor, but other schemes (e.g., weighted round-robin) are not excluded. In each run loop of the processor 4 (i.e., each cycle of the processor clock that times the program counter), the fetch stage 14 fetches a single instruction or a small "instruction bundle" (e.g., a two-instruction bundle or a four-instruction bundle), depending on the implementation. Then, depending on whether the instruction (based on its opcode) is a memory access instruction, an integer arithmetic instruction, or a floating-point arithmetic instruction, via the decode stage 16, each instruction is respectively issued to one of the LSU 55 or IALU 56 of the main run unit 18M or the FPU of the auxiliary run unit 18A. The LSU 55 and IALU 56 of the main run unit 18M run their instructions using the registers from the MRF 26M, and the specific registers within the MRF 26M are specified by the operands of the instruction. The FPU of the auxiliary run unit 18A performs operations using the registers in the ARF 26A and WRF 26W, where the specific registers within the ARF are specified by the operands of the instruction. In the example, the registers in the WRF can be implicit in the instruction type (i.e., predetermined for that instruction type). The auxiliary run unit 18A may also include circuitry in the form of auxiliary logic latches inside the auxiliary run unit 18A for storing certain internal states 57 for performing operations of one or more floating-point arithmetic instruction types.

[0059] In an example of fetching and running the instructions in a bundle, the individual instructions in a given instruction bundle run in parallel simultaneously along independent pipelines 18M, 18A (as Figure 1A shown). In an example of running a two-instruction bundle, the two instructions can run simultaneously along the corresponding auxiliary pipeline and the main pipeline. In this case, the main pipeline is arranged to run instruction types that use the MRF, and the auxiliary pipeline is used to run instruction types that use the ARF. Pairing the instructions into suitable complementary bundles can be handled by the compiler.

[0060] Each worker thread context has an instance of its own master register file (MRF) 26M and auxiliary register file (ARF) 26A (i.e., each barrel-threaded slot has an MRF and an ARF). The functions described herein with respect to the MRF or ARF are to be understood as operating on a per-context basis. However, a single shared weight register file (WRF) is shared among the threads. Each thread can access only the MRF and ARF of its own context 26. However, all currently running worker threads can access the common WRF. Thus, the WRF provides a common set of weights for use by all worker threads. In the example, only the supervisor can write to the WRF, and the workers can only read from the WRF.

[0061] The instruction set of processor 4 includes at least one type of load instruction whose opcode, when run, causes the LSU 55 to load data from the data memory 22 into the corresponding ARF 26A of the thread in which the load instruction is run. The location of the destination within the ARF is specified by the operand of the load instruction. Another operand of the load instruction specifies an address register in the corresponding MRF 26M that holds a pointer to the address in the data memory 22 from which the data is loaded. The instruction set of processor 4 also includes at least one type of store instruction whose opcode, when run, causes the LSU 55 to store data from the corresponding ARF of the thread in which the store instruction is run to the data memory 22. The location of the store source within the ARF is specified by the operand of the store instruction. Another operand of the store instruction specifies an address register in the MRF that holds a pointer to the address in the data memory 22 where the data is to be stored. Typically, the instruction set can include separate load instruction types and store instruction types, and / or at least one load-store instruction type that combines the load and store operations in a single instruction.

[0062] In response to the opcode of a relevant type of arithmetic instruction, an arithmetic unit (e.g., an FPU) in the auxiliary execution unit 18A performs an arithmetic operation as specified by the opcode, which includes operating on the values in the source registers specified in the corresponding ARF of the thread and optionally the source registers in the WRF. It also outputs the result of the arithmetic operation to the destination register in the corresponding ARF of the thread explicitly specified by the destination operand of the arithmetic instruction.

[0063] The processor 4 may also include a switching interface 51 for exchanging data between the memory 11 and one or more other resources, such as other instances of the processor and / or external devices such as a network interface or a network-attached storage (NAS) device. As described above, in an example, the processor 4 may form one of an array of interconnected processor tiles, each tile running a part of a broader program. Thus, the individual processors 4 (tiles) form part of a wider processor or processing system 6. The tiles 4 may be connected together via an interconnect subsystem to which they are connected via their respective switching interfaces 51. The tiles 4 may be implemented on the same chip (i.e., die) or different chips or a combination (i.e., the array may be formed of multiple chips, each chip including multiple tiles 4). Thus, the interconnect system and the switching interface 51 may correspondingly include internal (on-chip) interconnect mechanisms and / or external (inter-chip) switching mechanisms.

[0064] Figure 2 A simple logic circuit 100 for performing data block rotation operations that do not fall within the scope of the claims is shown.

[0065] The logic circuit 100 is configured to rotate input data blocks of size 4, 8, or 16 bytes. The logic circuit has a 16-byte input data array 101. In the case where a 4-byte block needs to be rotated, a subset 101a of the input data array 101 receives the 4-byte block. In the case where an 8-byte block needs to be rotated, the subset 101a is used for 4 of the 8 bytes, while another subset 101b is used for the remaining 4 bytes.

[0066] The logic circuit includes a 16:1 multiplexer array 102, an 8:1 multiplexer array 103, and a 4:1 multiplexer array 104. Each of the multiplexer arrays receives a block of a respective size S, and a control signal (not shown) indicating the amount of rotation. The amount of rotation indicates the number of bytes by which the block is to be rotated.

[0067] Each of the arrays 102 - 104 includes a series of S multiplexers, where the input j of each multiplexer i is connected to the byte (i + j) mod S of the input data array. Thus, the 0th input of the 0th multiplexer of the 4:1 array 104 is connected to byte (0 + 0) % 4 = 0 of the input data array, the 1st input is connected to (0 + 1) % 4 = 1, and so on. The 0th input of the 1st multiplexer of the 4:1 array 104 is connected to byte (1 + 0) % 4 = 1, the 1st input of the 1st multiplexer is connected to byte (1 + 1) % 4 = 2, and so on for all inputs of all multiplexers in the array. In this context, both "%" and "mod" are shorthand for the modulo operator. Thus, by arranging the inputs in this way, control signals can be applied to each multiplexer in the array to select the relevant bytes from the input data array to effect the desired rotation.

[0068] In this document, in some contexts, such as when discussing data arrays or multiplexer arrays, a convention of indexing starting from 0 is followed. Thus, the 0th element of an array can be the element that first appears in the array, the 1st element can be the second - appearing element, and so on. In other contexts, terms such as first, second, third, etc. can simply be used as labels to distinguish similar elements. The meaning intended will be apparent from the relevant context.

[0069] The result multiplexer array 105 then selects results from the relevant S:1 arrays 102 - 104 to provide result 106 again based on another suitable control signal, called the select signal. For example, the lowest log2(S) bits of the rotation amount can be used as the select signal.

[0070] Although this method is simple, using large multiplexers for large block sizes may be too slow to meet the timing constraints of relatively short processor clock cycles. In addition, a large amount of long wiring is required to connect the multiplexer inputs, resulting in routing congestion. It should be understood that the difficulty in meeting the timing constraints may be affected to some extent by the amount of logic before and / or after circuit 100 and the available timing margin, as well as other constraints related to the placement of circuit 100 within the processing unit and routing congestion within the processing unit. In addition, the wiring of each of the initial multiplexer arrays 102 - 104 only supports a single fixed block size.

[0071] Figure 3 An improved logic circuit 20 according to an example of the present disclosure is shown. The logic circuit 20 includes a circuit 200 configured to perform rotation operations for each of 16 - byte, 8 - byte, and 4 - byte blocks.

[0072] The logic circuit includes an input data array 201 configured to receive 16 - byte, 8 - byte, or 4 - byte data blocks to be rotated. Each input byte of the input data array 201 can be considered a data "channel".

[0073] The circuit 200 is arranged in a plurality of processing levels 202, 203, 204. Each level is configured to compute rotations related to successively larger block sizes (4, 8, 16). The first level 202 operates on the input data array 201. The subsequent levels 203, 204 operate on the output from the previous level as partial results.

[0074] Thus, the rotation operations for larger block sizes are effectively decomposed across multiple processing levels 202 - 204. It should be understood that the "level" in this context is a construct for discussing the logical arrangement of the components of circuit 200 and does not imply any particular physical layout of circuit 200.

[0075] First processing level 202 comprises four rotators 210a-210d.Each rotator 210 is configured to calculate the rotation of 4 byte blocks.Therefore, each rotator 210 is connected to the different corresponding 4 bytes of input data array 201.For example, rotator 210a is connected to the byte 0-3 of array 201, and rotator 210b is connected to byte 4-7, and the rest may be deduced by analogy.Each rotator 210 offers its output to the corresponding byte of first intermediate result array 205.That is to say, rotator 210a is connected to the byte 0-3 of first intermediate result array 205, and rotator 211b is connected to byte 4-7, and the rest may be deduced by analogy.Therefore, first processing level 202 comprises four parallel 4 byte rotators effectively, and each rotator operates on 4 data channels of input array.

[0076] Figure 4 The construction and operation of the rotator 210a of the first processing level 202 are shown in detail. The rotator 210a takes the form of a multiplexer array comprising four multiplexers 211a-211d. Each multiplexer 211 has four inputs connected to four corresponding bytes forming a subset of the input data array 201 on which the rotator 210a operates.

[0077] The multiplexer array is similar to the one above. Figure 2 The 4:1 array discussed above operates in the same manner, with input j of each multiplexer i connected to byte (i+j) mod 4 of the corresponding 4-byte subset of input data array 201 connected to rotator 210a. Thus, the 0th input of multiplexer 211a (i.e., the 0th multiplexer) is connected to byte (0+0)%4=0 of the input data array, the 1st input is connected to (0+1)%4=1, and so on. The 0th input of multiplexer 211b (i.e., the 1st multiplexer) is connected to byte (1+0)%4=1, the 1st input of the 1st multiplexer is connected to byte (1+1)%4=2, and so on for all inputs of all multiplexers in the array.

[0078] The output of each multiplexer 211 in the array is provided to a corresponding byte of the first intermediate result array 205. That is, the 0th multiplexer 211 is connected to byte 0 of the intermediate result array 205, the 1st multiplexer 211 is connected to byte 1 of the intermediate result array 205, and so on.

[0079] The rotator 210a is configured to receive a control signal 212 indicating a desired amount of rotation. For example, the control signal 212 may include a 2-bit signal representing an amount of rotation from 0 to 3 (i.e., up to block size - 1). The control signal 212 is provided to each multiplexer 211, which selects the input corresponding to the amount of rotation. This causes each multiplexer in the array to select the relevant bytes from the input data array to effect the desired rotation. If the amount of rotation is 0, the multiplexer array acts as a pass-through.

[0080] The rotators 210b - 210d operate substantially the same, except for receiving inputs from and providing outputs to different bytes of the input array 201 and the intermediate result array 205.

[0081] In terms of expressing the relationship between a given multiplexer 211 of any rotator 210 and the complete input data array 201, the input j of the i-th S:1 multiplexer should be connected to the subsequent byte of the input data array 201:

[0082] (floor(i / S) * S) + (j % S)

[0083] In the above formula, i is the index of the multiplexer 211 in the entire first processing level 202, such that the multiplexers 211a - 211d of the rotator 210a are indexed 0 - 3, the multiplexers 211 of the rotator 210b are indexed 4 - 7, and so on.

[0084] Returning to Figure 3 , the second processing level 203 includes two 8-byte rotators 220a, 220b. Each 8-byte rotator is connected to a different 8-byte of the first intermediate result array 205. For example, rotator 220a receives inputs from bytes 0 - 7 of array 205, and rotator 220b receives inputs from bytes 8 - 15 of array 205. Each rotator 220a outputs a rotated 8-byte block to the second intermediate result array 215. Thus, rotator 220a generates a rotated 8-byte block (i.e., rotated relative to the corresponding 8 bytes of the input array 201) from the two rotated 4-byte blocks output by the first level 202.

[0085] Figure 5 The rotator 220a of the second processing level 203 is shown in more detail. The rotator takes the form of an array of eight multiplexers 221a - 221h.

[0086] Each multiplexer is a 2:1 multiplexer. Input 0 of the i-th multiplexer is connected to byte i mod 4 of array 205. Input 1 of the i-th multiplexer is connected to byte (i mod 4) + 4 of array 205. Thus, each multiplexer 221 effectively selects between the corresponding bytes of the first rotated block (i.e., bytes 0 - 3 output from rotator 210a to array 205) or the second rotated block (i.e., bytes 4 - 7 output from rotator 210b to array 205), where input 0 is connected to the byte of the first rotated block and input 1 is connected to the byte of the second rotated block.

[0087] The output of each multiplexer 221 in the array is provided to the corresponding byte of the second intermediate result array 215. That is, the 0-th multiplexer 221 is connected to byte 0 of the intermediate result array 215, the 1-st multiplexer 221 is connected to byte 1 of the intermediate result array 215, and so on.

[0088] Rotator 220a is configured to receive control signal 222. The control signal 222 takes the form of an 8-bit bitmask 0xF0 (i.e., 11110000) and rotates by an amount corresponding to the desired rotation amount. The desired rotation amount is a value between 0 - 7. Thus, for a rotation amount of 1, the bitmask would be 11100001, for a rotation amount of 2, the bitmask would be 11000011, and so on.

[0089] The bitmask is split by splitter 223 such that each bit of the bitmask of control signal 222 is directed to the corresponding multiplexer 221. In other words, bit 0 of the bitmask is provided to the 0-th multiplexer 221a, bit 1 of the bitmask is provided to the 1-st multiplexer 221b, and so on. The splitter 223 can include suitable circuitry and can, for example, include wiring to carry each of the bits to its corresponding multiplexer 221. In some examples, the splitter 223 can be embodied by just the wiring.

[0090] Thus, rotator 220a takes two 4-byte blocks that have been rotated by any amount and outputs an 8-byte block that has been rotated by the same arbitrary amount. If the rotation amount is 0 (i.e., the bitmask is 11110000), then rotator 220a acts as a pass-through.

[0091] Rotator 220b operates in substantially the same manner except that it receives inputs from different bytes of the first intermediate result array 201 and the second intermediate result array 215 and provides outputs to them.

[0092] Back to Figure 3, the third processing level includes a 16 - byte rotator 230. The rotator 230 takes all 16 bytes of the second intermediate result array 215 as input, and all 16 bytes include two rotated 8 - byte blocks output by the rotator 220. The rotator 230 outputs the rotated 16 - byte block to the final result array 225.

[0093] Figure 6 Rotator 230 is shown in more detail. Again, rotator 230 takes the form of a multiplexer array, in this case including 16 2:1 multiplexers 231a - 231p. To improve the clarity of the drawing, some of the multiplexers 231 are omitted.

[0094] The construction and operation of rotator 230 are similar to rotator 220, although it is adapted to operate on 8 - byte input blocks rather than 4 - byte input blocks. Thus, input 0 of the i - th multiplexer is connected to byte I mod 8 of array 215. Input 1 of the i - th multiplexer is connected to byte (i mod 8)+8 of array 215. Thus, each multiplexer 231 effectively selects between the corresponding byte of the first rotated block (i.e., bytes 0 - 7 output from rotator 220a to array 215) or the second rotated block (i.e., bytes 8 - 15 output from rotator 220b to array 215), where input 0 is connected to the byte of the first rotated block and input 1 is connected to the byte of the second rotated block.

[0095] Similar to rotator 220, rotator 230 is configured to receive control signals in the form of bit - mask rotations. However, in this case, the bit - mask is a 16 - bit bit - mask 0xFF00. As mentioned before, the rotation of the bit - mask reflects the desired amount of rotation. The received bit - mask is split by a splitter 233, which distributes the bits of the bit - mask to their respective multiplexers 230.

[0096] In some examples, circuit 200 may include pipeline registers 206 disposed between at least some of the processing levels 202. These are shown in Figure 8 This may help meet the timing requirements.

[0097] Figure 7 It is shown that the logic circuit 20 may also include a control signal generator 250, which includes circuitry configured to generate control signals 212, 222, 232 to control the rotators to perform the desired rotations. To generate such signals, the control signal generator may be provided with the block size 251 to be rotated and the desired amount of rotation 252.

[0098] As discussed in more detail below, these can form part of the rotate instruction 253, or can form part of another instruction (such as a pack or extract instruction) that involves a rotate operation. That is, the rotation amount 252 and the block size 251 can be indicated by the instruction. For example, either or both of the rotation amount 252 and the block size 251 can be an operand or indicated by the opcode. In other examples, the rotation amount 252 and / or the block size 251 can be read from one or more registers as part of the execution of the instruction. It can also be the case that the instruction indicates a value that can compute the block size or the rotation amount.

[0099] As noted above, the control signal 212 for controlling the rotator 210 at the first level 202 is the desired rotation amount 252. For a block rotation of a rotation amount 252 greater than 3 for a block of size 251 greater than 4 bytes, the signal 212 represents the desired rotation amount mod 4.

[0100] As noted above, the control signal 222 for controlling the rotator 220 at the second level 203 is the bit mask 0xF0 rotated by the rotation amount 252. Thus, the control signal generator 250 can include suitable circuitry for rotating the bit mask 0xF0 by the rotation amount 252. For example, the control signal generator 250 can include (or otherwise access) a look-up table that stores the rotated version of the bit mask. The control signal generator 250 can then select the stored rotated bit mask based on the desired rotation amount 252. In another example, the control signal generator 250 can include circuitry for performing shift operations to rotate the bit mask by the desired amount. For a block rotation of a rotation amount 252 greater than 7 for a block of size 251 greater than 8 bytes, the bit mask is rotated by the rotation amount 252 mod 8.

[0101] As further pointed out above, the control signal 232 for controlling the third level rotator 230 is the bit mask 0xFF00 rotated by the rotation amount 252. Thus, the control signal generator 250 can also include circuitry similar to that described above for the rotation of the bit mask 0xF0 for performing this rotation, such as a suitable look-up table or shift circuitry.

[0102] To further facilitate the understanding of the operation of the logic circuit 20, an example of the circuit 20 in use will now be discussed.

[0103] To rotate a 4-byte block, the 4-byte block is provided as input to bytes 0-3 of the input data array 205. The control signal generator 250 generates the control signal 212 corresponding to the desired rotation amount, and this control signal 212 is provided to the rotator 210a. The rotator 210a performs the rotation, and the result of the rotation is output to bytes 0-3 of the intermediate result array 205.

[0104] Since only 4-byte block rotation is required, the control signal generator 250 generates unrotated bitmasks as control signals 222 and 232 for levels 203 and 204, causing them to act as throughputs. Thus, the rotated blocks in bytes 0-3 of the intermediate result array 201 pass through levels 203 and 204 and are output to bytes 0-3 of the output array 225. Generally, when one level acts as a throughput, subsequent levels will also act as throughputs.

[0105] If needed, the logic circuit 20 can input corresponding 4-byte blocks into the corresponding rotators 210 via the relevant bytes of the input data array 201 to perform rotation of up to four 4-byte blocks in parallel.

[0106] Similarly, if fewer than four 4-byte blocks are rotated in parallel, clock gating can be used to disable the circuits not used for the desired rotation. For example, if only one 4-byte block is rotated, rotators 210b-210d and 220b can be disabled. The control signal generator 250 can correspondingly generate appropriate control signals 242 for enabling or disabling the clock gating 207 that may be included in the circuit 200 (see Figure 8 )

[0107] To rotate an 8-byte block, the 8-byte block is provided as an input to bytes 0-7 of the input data array 205. The control signal generator 250 generates a control signal 212 based on the desired rotation amount according to the instruction, and this control signal 212 is provided to rotators 210a and 210b. Rotators 210a and 210b perform the rotation, and the results of the rotation are output to bytes 0-3 and 4-7 of the first intermediate result array 205 respectively.

[0108] The control signal generator 250 also generates a control signal 222 based on the desired rotation amount, and this control signal 222 is provided to rotator 220a. Rotator 220a outputs the rotated 8-byte block to bytes 0-7 of the second intermediate result array 215.

[0109] The control signal generator 250 also generates a control signal 232 for level 204, and this control signal 232 causes rotator 230 to act as a throughput. Thus, the rotated 8-byte block is output to bytes 0-7 of the output array 225.

[0110] In a similar manner to that discussed above for 4-byte blocks, two 8-byte blocks can be rotated in parallel using the logic circuit 20. Additionally, the control signal 242 can provide clock gating control signals 242 to disable the circuits not used for rotating a single 8-byte block.

[0111] To rotate a 16-byte block, the 16-byte block is provided as input to data array 205. Control signal generator 250 generates control signal 212 based on the desired amount of rotation, and this control signal 212 is provided to rotators 210a - 210d. Rotators 210a - 210d perform the rotation of their respective 4-byte blocks, and the results of the rotation are output to bytes 0 - 3, 4 - 7, 8 - 11, and 12 - 15 of the first intermediate result array 205, respectively.

[0112] Control signal generator 250 also generates control signal 222 based on the desired amount of rotation, and this control signal 222 is provided to rotators 220a and 220b. Rotator 220a outputs the rotated 8-byte block to bytes 0 - 7 of the second intermediate result array 215. Rotator 220b outputs the rotated 8-byte block to bytes 7 - 15 of the second intermediate result array 215.

[0113] Control signal generator 250 also generates control signal 232 based on the desired amount of rotation, and this control signal 232 is provided to rotator 230. Rotator 230 outputs the rotated 16-byte block to output array 225.

[0114] The examples discussed above can be incorporated into the processing unit discussed above with respect to Figure 1A For example, the execution unit 18 of processor 4 can include logic circuit 20. Circuit 20 can be implemented, for example, in the FPU 18A of processor 4.

[0115] More specifically, execution unit 18 can receive an instruction (i.e., computer program instruction) to execute a rotated data block. In one example, the instruction is a rotate instruction, which has the sole purpose of rotating a data block. However, it can also be the case that another instruction is provided, which involves rotating a data block along with additional processing of the rotated data block. That is, the instruction can include multiple operations, including a rotate operation to be performed by the logic circuit.

[0116] As described above, the received instruction 253 indicates the block size 251 and the desired amount of rotation 252 (e.g., as an operand, as part of the opcode, or read from a register as part of the execution of the instruction). Although not shown in Figure 6 the instruction 253 can also include an indication of the memory location of the block to be rotated. This can take the form of an address or a pointer to a register pointing to a storage address. In other examples, the memory address storing the block to be rotated can be implicit. Instruction 253 can optionally include an indication of the memory location for storing the rotated block. Similarly, this can include an address or a pointer to a register pointing to a storage address. Likewise, it can be implicit in the instruction.

[0117] Upon receiving the received instruction, the run unit 18 loads the data block into the input data array 201. For example, the run unit 18 may receive an instruction as discussed herein with respect to Figure 1A the instruction discussed. Specifically, the fetch stage 14 may fetch the instruction from the instruction memory 11 and then issue the instruction to the run unit 18 via the decode stage 16.

[0118] Based on the required amount of rotation 252, the control signal generator 250 generates appropriate control signals 212, 222, 232 to control the circuit 200. The data block is then processed by the circuit 200 and the output is read from the output array 225. If the instruction includes a memory location for storing the rotated block, the rotated block can be stored at that memory location.

[0119] Figure 8 The inclusion of the clock gating 207 and the pipeline register 206 is shown in more detail. For clarity, the input array and the output array are omitted in this figure.

[0120] The pipeline register 206 is provided between consecutive processing levels. For example, the first pipeline register 206-1 is located between the first processing level 202 and the second processing level 203. The second pipeline register 206-2 is located between the second processing level 203 and the third processing level 204. It should be understood that in some implementations, the pipeline register 206 may not be required, or only one pipeline register 206 may be required. By providing a hierarchical structure including levels 202, 203, 204, it becomes simpler to insert the pipeline register 206 to retime the logic circuit 20. In addition, the pipeline register 206 is shared across operations applied to different block sizes.

[0121] In some examples, one or both of the intermediate result arrays 205, 215 may include the pipeline register 206. That is, in some cases, each of the intermediate result arrays 205 and / or 215 is the pipeline register 206. However, in other examples, one or both of the intermediate result arrays 205, 215 may be wiring. As discussed herein, whether the arrays 205 and 215 are pipeline registers or wiring may depend on the timing requirements.

[0122] Figure 8Further shown are a plurality of clock gating 207. The clock gating 207 in this example is configured to disable one or more data path channels of the logic circuit 20. For example, each gating 207 can be arranged to disable four channels. In the illustrated example, the first clock gating 207-1 can disable channels 0-3, the second clock gating 207-2 can disable channels 4-7, the third clock gating 207-3 can disable channels 8-11, and the fourth clock gating 207-4 can disable channels 11-15.

[0123] Thus, in the case where only one 4-byte block is to be rotated, the clock gating 207-2, 207-3, and 207-4 can disable their corresponding channels. Similarly, if an 8-byte block is to be rotated, then the clock gating 207-3 and 207-4 can disable their corresponding channels. In the case where the multiplexer disables some of its channels, the remaining channels can be used as throughputs.

[0124] It will be appreciated that this is an example of the configuration of the clock gating 207. In other examples, the gating 207 can disable more or fewer channels (e.g., 2 or 8 channels). Additionally, in some examples, the gating 207 can be arranged to disable a particular multiplexer.

[0125] Now will discuss Figure 9 and Figure 10 , which respectively show a first code example and a second code example of an instruction including a rotation operation.

[0126] The rotation operation discussed above can find particular utility in a class of instructions (referred to herein as data movement instructions). Data movement instructions can be used in the processor 4 to accelerate the movement of any aligned data.

[0127] For example, it may be the case that the architecture of the execution unit imposes certain constraints on accessing data in memory. Assume that the execution unit is configured to run load instructions for data units that are four bytes wide. In this case, each load instruction may only load data from a memory address that represents a 4-byte fine partition of the memory. For example, if the starting address of the memory is 0x80000, a load instruction may cause four bytes of data to be loaded starting from memory address 0x80000, or may be used to cause four bytes of data to be loaded starting from memory address 0x80004. However, given the architectural constraints of the given processing unit, it is not possible to load data starting from memory address 0x8002 because that memory address is not aligned with the size of the memory access. Similar constraints may apply to store operations. Such constraints may lead to problems in, for example, a "worker" program (i.e., a program run by a worker thread), which may be used jointly to process a large batch of application data (e.g., in the training of a machine learning model). Data movement instructions can be used to align misaligned data and thus accelerate the processing of the application data.

[0128] Although discussed in more detail above with respect to FIG. 1, relevant aspects of an example processing pipeline are summarized below. Data movement instructions may be available in a "worker" program that is used jointly to process a large batch of application data. The worker program is run sequentially by a tile processor (e.g., processor 4 of FIG. 1) through a fixed-length pipeline. Pipeline latency is hidden by keeping multiple worker threads resident in each tile 4 and scheduling these worker threads in a barrel-thread manner over multiple clock cycles.

[0129] In a first pipeline stage, an instruction fetch stage 14 fetches the raw instruction word from a runnable region of a tile-coupled memory 11 into a local buffer 53 inside the tile 4.

[0130] In a second pipeline stage, decoding logic (e.g., decoding stage 16) converts the fetched instruction word into an internal data construct that describes how the rest of the pipeline must be controlled to execute the instruction. For a data movement instruction, fields of this data structure will signal that the FPU of the auxiliary execution unit 18A (and the data movement pipeline therein) is enabled, describe the operation to be performed via an opcode, and provide source and destination operand addresses.

[0131] In a third stage, operands are read. For an instruction that uses the logic circuit 20 discussed herein, a series of registers in the ARF 26A are indexed according to the operand addresses decoded by the instruction. The read operand data is provided to the data movement pipeline inside the FPU.

[0132] In the next several stages, the data movement pipeline executes instructions to process data blocks (operand data). Data movement instructions are, for example, "pack" or "extract" operations, having a block size of, for example, 4, 8, or 16 bytes. The data blocks are data operands read from the ARF26A. The "pack" and "extract" operations can also be controlled by each worker status register, as described below.

[0133] After running the instruction, the output is written back to the ARF 26A at the location specified in the operands of the instruction.

[0134] Figure 9 A code example of executing a pack instruction is shown. The pack instruction uses wrapping to copy an arbitrary sequence of consecutive bytes from a first block to an arbitrary location in a second block. The logic circuit 200 is used to align the bytes of the first block to the correct output location. The rotation amount is derived from a field in the $PACK worker status register and is implicitly used by the pack instruction. Depending on the configuration of the $PACK worker status register, a mask is used to select any number of consecutive bytes from the second data block or the rotated first data block.

[0135] Specifically, reference numeral 91 indicates loading an 8-byte first data block into the memory location $a0:1. Two ldconst instructions are used to load each 4-byte of the block into $a0 and $a1. That is, the first operand of the ldconst instruction indicates the location where the data is loaded, and the second operand indicates the data. In this example, the ldconst instruction is used to load the location $a0:1 with pseudo data. In practice, it should be understood that the memory location will be loaded with application data.

[0136] Reference numeral 92 indicates the corresponding instruction for loading the second data block into the location $a2:3.

[0137] Subsequently, the pack instruction is configured. In the example, the pack instruction is intended to insert 3 bytes from the first block at location 1 into the second block at location 5. For this purpose, the register $m0 is set with a value reflecting the desired configuration, indicated by reference numeral 93. In the example, the setzi instruction is the instruction configured to set the register. This instruction takes a first operand indicating the register ($m0) and a second operand including the value to be stored in the register.

[0138] In the example, the two least significant bytes (0x03) indicate the number of bytes to be inserted. The next two bytes (0x01) indicate the location in the first block from which those bytes will be read. The last byte (0x5) indicates the location in the second block where the bytes will be inserted.

[0139] Then, in step 94, the value is copied from register $m0 to the $PACK worker status register using the put instruction.

[0140] Finally, the pack instruction is run, as shown by reference numeral 95. The first operand indicates where the result of the instruction will be output, where the second and third operands indicate the positions of the first and second blocks in memory, respectively. Thus, in the example, the result is written back to $a0:1, which is where the first block was read.

[0141] The opcode of the pack instruction indicates the block size. In this example, the opcode is "pack64", indicating an 8-byte (i.e., 64-bit) block size. Different opcodes can be provided for different-sized blocks, such as pack128 for 16-byte blocks and pack32 for 4-byte blocks.

[0142] When the pack instruction is run, the execution unit determines the amount of rotation required based on the value in $pack, and determines the block size to be rotated based on the opcode. Thus, the control signal generator 250 generates appropriate control signals to configure the logic circuit to perform the required rotation. In the example discussed above, the block size is 8, and the bytes of the first block are rotated by four bytes to match the output position.

[0143] Figure 10 A code example for executing the extract instruction is shown. The extract instruction uses wrapping to extract a continuously positioned block of bytes from the concatenation of two input blocks. The $EXTRACT worker status register indicates the "pivot point", representing the starting index for selecting bytes from the double-length block formed by the concatenation. A mask is used to select the relevant bytes from either input block. The resulting bytes are then rotated to the correct output position using the logic circuit 200.

[0144] Specifically, reference numeral 1001 indicates loading a 16-byte first data block into memory location $a0:3. Four ldconst instructions are used to load each 4-byte portion of the block into $a0 to $a3, respectively. Reference numeral 1002 indicates the corresponding instruction for loading the second data block into location $a4:7.

[0145] Reference numerals 1003 and 1004 indicate setting the $EXTRACT register with the value 0x7, reflecting the fact that the output starts being extracted at byte 7 of the concatenation of the two input blocks.

[0146] Reference numeral 1005 shows the execution of the extract operation. Similar to the pack instruction, the first operand indicates where the result of the instruction will be output, while the second and third operands indicate the positions of the first and second blocks in memory, respectively. Thus, in the example, the result is written back to $a0:3, which is where the first block was read.

[0147] The opcode of the extract instruction also indicates the block size. In this example, the opcode is "extract128", indicating a 16-byte (i.e., 128-bit) block size. Different opcodes can be provided for blocks of different sizes, such as extract64 for 8-byte blocks and extract32 for 4-byte blocks.

[0148] When the extract instruction is run, the execution unit determines the amount of rotation required based on the value in $EXTRACT and determines the block size to be rotated based on the opcode. Thus, the control signal generator 250 generates appropriate control signals to configure the logic circuit to perform the required rotation.

[0149] Both the packing and extract instructions are discussed in more detail in co-pending U.S. Patent Application No. 18 / 053948, the content of which is incorporated herein by reference in its entirety.

[0150] A further discussion of the multi-tile processing unit now follows. As discussed above, the processor 4 can form part of a multi-tile processing device. There are many possible different manifestations of suitable processing devices, which can take the form of a chip. Graphcore has developed an intelligent processing unit (IPU), which is described, for example, in U.S. Patent Application Nos.: US 2019 / 0121387A1; US 2019 / 0121388A1; US 2019 / 0121777A1; US 2020 / 0319861A1, the content of which is incorporated herein by reference. Figure 1BIt is a height schematic diagram of the IPU. The IPU includes multiple tiles 1103 on a silicon die, and each tile includes a processing unit with local memory (e.g., the processing unit 4 discussed above). The tiles communicate with each other using time-deterministic switching. The switching fabric 1101 (sometimes referred to as a switch or switching fabric) is connected to each tile via corresponding output wiring and can be connected to each tile via its corresponding input wiring through switching circuits controllable by each tile. A synchronization module (not shown) is operable to generate synchronization signals to switch between a computing phase and a switching phase. The tiles run their local programs during the computing phase according to a common clock, which can be generated on the die or received by the die. At a predetermined time in the switching phase, a tile can run a send instruction from its local program to send a data packet onto the output set of its connected wiring. The data packet is destined for at least one receiving tile, but does not have a destination identifier identifying that receiving tile. At a predetermined switching time, the receiving tile runs a switching control instruction from its local program to control the switching circuit to connect the input set of its wiring to the switching fabric to receive the data packet at the receiving time. The send time (at which the data packet is scheduled to be sent from the sending tile) and the predetermined switching time are managed by the common clock relative to the synchronization signal.

[0151] Time-deterministic switching allows for efficient transfer between the tiles on the die. Each tile has its own local memory, which provides data storage and instruction storage. As described herein, the IPU is additionally connected to an external memory from which data can be transferred onto the IPU for use by the tiles via the fabric chip.

[0152] The tiles 1103 of the IPU can be programmed such that a data packet sent by a send instruction from its local program is intended to access memory (a memory access packet) or has another IPU connected in a cluster or system at its destination. In those cases, the data packet is sent by the originating tile 1103 onto the switching fabric but is not picked up by a receiving tile within the IPU. Instead, the switching fabric causes the tile to be provided to a suitable connector C1, C2, etc. for external communication from the IPU. Packets for off-chip communication are generated to include information defining their ultimate off-chip destination rather than the external port from which the packet is to be sent. When compiling code for a tile, the principle of time-deterministic switching can be used to send a packet to an external port to identify the external port of the packet. For example, a memory access packet can identify a memory address. A packet intended for another IPU can include the identifier of the other IPU. This information is used by the routing logic on the fabric chip to correctly route the off-chip packets generated by the IPU.

[0153] Figure 1BThe illustration therein shows five example regions of an IPU chip of an example, which are separated by four boundaries 1105, represented by dashed lines. Note that the dashed lines represent the abstract boundaries 1105 of the abstract regions on the processor chip, shown for illustrative purposes; the boundaries 1105 do not necessarily represent physical boundaries on the IPU chip.

[0154] In addition to incorporating logic circuits into processing units in the form of tile processors, it should be understood that logic circuits can be incorporated into various processing units or devices.

[0155] Within the scope of the present disclosure, various modifications can be made to the examples discussed herein.

[0156] For example, it should be understood that the examples shown above can be easily extended to larger block sizes that are powers of two. For example, to rotate a 32-byte block, the existing circuitry is replicated (i.e., replicated horizontally such that each layer 202 includes twice the number of rotators 210, 220, 230), and the input / output / intermediate array includes 32 bytes. Then an additional layer including 32 2:1 multiplexers can be added, which are controlled by a bitmask that is twice the width of the previous layer (i.e., 0xFFFF0000). This pattern can be replicated to provide circuitry for rotating blocks of any size that are powers of two. Similarly, layer 204 can be omitted to provide circuitry for rotating only 4-byte and 8-byte blocks.

[0157] Advantageously, the examples discussed herein provide a means for rotating blocks of a relatively large size S, which avoids complex S:1 arrays. Thus, smaller and faster multiplexers are employed, and the need for routing a large number of long data paths is avoided. This helps to keep the timing paths short.

[0158] In addition, the examples discussed herein help to minimize the area cost of processor features to provide the best performance per unit area. The examples involve hardware resources where rotation instructions are reused across all sizes, rather than requiring different hardware resources for different block sizes. The techniques discussed here naturally break down into levels and thus can be easily broken down into pipeline stages by inserting appropriate pipeline registers.

[0159] In addition, these examples allow for trivial clock gating of unused hardware resources such that only the required hardware resources are enabled, thereby minimizing power consumption.

Claims

1. An operating unit configured to execute byte-by-byte rotation operations on an input data block by running computer program instructions, the operating unit comprising: Logic circuitry, comprising: An input data array for receiving an input data block comprising N bytes; Two first-layer multiplexer arrays, each first-layer multiplexer array being configured to: Receive a first-layer data block, the first-layer data block comprising a corresponding byte subset of the input data block; Receive first-layer control signals; Rotate the first-layer data block by an amount indicated by the first-layer control signals; The two first-layer multiplexer arrays being configured to output a first rotated first-layer data block and a second rotated first-layer data block respectively; A second-layer multiplexer array configured to receive second control signals, the second-layer multiplexer array comprising N multiplexers, each multiplexer being configured to select between corresponding bytes of the first rotated first-layer data block and the second rotated first-layer data block based on the second control signals to output a rotated second-layer data block, and A control signal generator configured to generate the first-layer control signals and the second-layer control signals based on the received computer program instructions.

2. The operating unit according to claim 1, wherein Each first-layer data block may comprise N / 2 bytes.

3. The operating unit according to claim 1 or 2, wherein: The input data array is configured to receive an input data block comprising M bytes, where M > N; The logic circuitry comprises: Four first-layer multiplexer arrays to output first to fourth rotated first-layer blocks; Two second-layer multiplexer arrays, the first of the second-layer multiplexer arrays being configured to receive the first rotated first-layer block and the second rotated first-layer block and output a first rotated second-layer data block, the second of the second-layer multiplexer arrays being configured to receive the third rotated first-layer block and the fourth rotated first-layer block and output a second rotated second-layer block; A third-layer multiplexer array configured to receive third-layer control signals, the third-layer multiplexer array comprising M multiplexers, the M multiplexers being configured to select between corresponding bytes of the first rotated second-layer data block and the second rotated second-layer data block based on the third control signals to output a rotated third-layer data block; and The control signal generator is configured to generate the third-layer control signals.

4. The operating unit according to any one of the preceding claims, wherein, The logic circuitry comprises an intermediate result array configured to receive the output of the first-layer multiplexer arrays.

5. The operating unit according to any one of the preceding claims, wherein, Each first-layer multiplexer array comprises a plurality of S:1 multiplexers.

6. The operating unit according to any one of the preceding claims, wherein, The second-layer control signals comprise an N-bit bit mask, each bit in the bit mask serving as a control signal for a corresponding one of the N multiplexers of the second multiplexer array.

7. The operating unit according to claim 6, wherein, The logic circuitry comprises a splitter for splitting the bit mask and providing corresponding bits to corresponding multiplexers.

8. The operating unit according to claim 6 or 7, wherein, The control signal generator is configured to rotate the bit mask, wherein the amount of rotation of the bit mask causes the output to rotate the rotated second layer data block by the same amount.

9. The operating unit according to claim 8, wherein, The control signal generator is configured to rotate the bit mask by selecting a stored rotated bit mask from a lookup table.

10. The operating unit according to claim 3, wherein, The third layer control signal is an M-bit bit mask.

11. The operating unit according to any one of the preceding claims, wherein, The control signal generator generates a second layer control signal that causes the second layer multiplexer array to function as a pass-through.

12. The operating unit according to any one of the preceding claims, wherein The logic circuit includes: A plurality of data path channels; One or more clock gating units configured to disable one or more of the data path channels, and The control signal generator is configured to generate clock gating control signals to control the one or more clock gating units.

13. The operating unit according to any one of the preceding claims, wherein, The logic circuit includes pipeline registers disposed between the first layer multiplexer array and the second layer multiplexer array.

14. The operating unit according to any one of the preceding claims, wherein, The computer program instructions are rotation instructions configured to rotate the input data block.

15. The operating unit according to any one of the preceding claims, wherein, The computer program instructions include a plurality of operations, wherein the rotation operation is one of the plurality of operations.

16. The operating unit according to claim 15, wherein, The computer program instructions include: A packing instruction configured to copy a sequence of consecutive bytes from a first location in a first data block to a second location in a second data block, or An extraction instruction configured to extract a sequence of consecutive bytes from the concatenation of a first data block and a second data block.

17. The operating unit according to any one of the preceding claims, wherein, The operating unit is configured to generate the first layer control signal, the second layer control signal, and optionally the third layer control signal based on values indicated by the computer program instructions.

18. The operating unit according to claim 17, wherein, The values indicated by the computer program instructions are values included in one or more of the following: One or more operands of the computer program instructions, The opcode of the computer program instructions, Values stored in one or more registers associated with the computer program instructions.

19. A processing unit, including the operating unit according to any one of the preceding claims.

20. A processing device, the processing device including a plurality of processing units, wherein, At least one of the processing units is defined according to claim 19.

21. The processing device according to claim 20, wherein, The processing unit is a tile processor that communicates via a switching fabric that implements time-deterministic switching.

22. A method implemented in an operating unit, the method including: Receiving an input data block including N bytes; Providing a first layer data block including a corresponding byte subset of the input data block to a first first layer multiplexer array and a second first layer multiplexer array; Providing first layer control signals to the first first layer multiplexer array and the second first layer multiplexer array; Rotating the first layer data block by an amount indicated by the control signal to output a first rotated first layer data block and a second rotated first layer data block; Providing the first rotated first layer data block and the second rotated first layer data block to a second layer multiplexer array, the second layer multiplexer array including N multiplexers, A second control signal is provided to the second layer multiplexer array to select between corresponding bytes of the first rotated first layer data block and the second rotated first layer data block based on the second control signal.

Citation Information

Patent Citations

  • Processing device for handling misaligned data

    US12124699B2

  • Synchronization in a multi-tile processing array

    US20190121387A1

  • Compiler method

    US20190121388A1

  • Instruction set

    US20190121777A1

  • Compiling a Program from a Graph

    US20200319861A1