Tile Processor Communication Structure

Through the block-based structure design, efficient communication between processor circuits is achieved, performance and power consumption problems when complexity increases are solved, and scaling and communication efficiency are supported in different configurations.

CN119404182BActive Publication Date: 2025-07-29APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202380048358.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-06-23
Filing Date
2023-05-11
Publication Date
2025-07-29
Estimated Expiration
2043-05-11

AI Technical Summary

Technical Problem

Existing processor circuit designs are difficult to meet the requirements of performance, power consumption and circuit area when complexity increases, especially in the design of component communication structures.

Method used

The tile-based structure design is adopted, which supports input and output configurations in different directions, and realizes efficient communication between inputs through single-cycle arbitration and multi-cycle priority update schemes.

Benefits of technology

It facilitates the scaling of processor design, improves communication efficiency and power management, and meets the performance requirements of different processor configurations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119404182B_ABST
    Figure CN119404182B_ABST
Patent Text Reader

Abstract

Techniques related to a processor communication structure are disclosed. In some embodiments, a processor includes a plurality of client circuits and a structure circuit including at least a first instance and a second instance of a tile. The tile may include: a client input configured to interface with the client circuits, a tile input configured to interface with one or more other tile instances, and communication resources that can be assigned to the client input and the tile input. The communication resources may include: a plurality of internal links, a client output configured to interface with the client circuits, and a tile output configured to interface with one or more other tile instances. A control circuit may assign, in a given cycle, the communication resources of a given tile instance to at least a portion of the client input and the tile input based on priority information for a next cycle. The control circuit may update the priority information based on the assignment result over a plurality of cycles.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art Technical Field

[0002] The present disclosure relates generally to processor architectures, and more particularly to a tiled structure for communication between processor circuits.

[0003] Description of Related Technologies

[0004] As computer processors increase in complexity, the design of communication structures between components can be important to meeting performance, power consumption, and circuit area targets. For example, in the context of graphics processors, the number of shader pipelines typically increases over time, and communication can occur between various agents (e.g., cache controllers, different types of execution pipelines, fixed-function circuits such as samplers, ray accelerator circuits, etc.). Additionally, traditional designs can be difficult to scale to meet the requirements of a given architecture. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Figure 1A is a diagram illustrating an overview of example graphics processing operations according to some embodiments.

[0006] Figure 1B is a block diagram illustrating an example graphics unit according to some embodiments.

[0007] Figure 2 is a block diagram illustrating an example client configured to communicate via a fabric according to some embodiments.

[0008] Figure 3 is a block diagram illustrating an example tile-based structure circuit according to some embodiments.

[0009] Figure 4 is a block diagram illustrating a detailed example tile including multiple slices according to some embodiments.

[0010] Figure 5 is a block diagram illustrating a detailed example slice according to some embodiments.

[0011] Figure 6 is a flow diagram illustrating an example single-cycle arbitration technique for tiles according to some embodiments.

[0012] Figure 7 is a diagram illustrating an example priority comparison for arbitration according to some embodiments.

[0013] Figure 8 is a diagram illustrating an example priority mask matrix for arbitration according to some embodiments.

[0014] Figure 9is a flowchart illustrating an example multi - cycle priority update technique according to some embodiments.

[0015] Figure 10 is a flowchart illustrating an example method according to some embodiments.

[0016] Figure 11 is a block diagram illustrating an example computing device according to some embodiments.

[0017] Figure 12 is an illustration showing an example application of the disclosed systems and devices according to some embodiments.

[0018] Figure 13 is a block diagram illustrating an example computer - readable medium storing circuit design information according to some embodiments. Detailed Description

[0019] In the disclosed embodiments, the tile - based fabric is configured to route communications between various processor agents. Examples of agents in a graphics context include memory interfaces, one or more cache controllers, shader data path circuits, fixed - function circuits, etc. In other contexts (e.g., central processing unit or microcontroller), similar fabrics can be implemented.

[0020] The disclosed tile - based fabric can have different configurations for different tiles, for example, in terms of the number of inputs and outputs in different directions and internal communication resources. Additionally, the fabric can be configured to have a different number of tiles in different designs. This can advantageously facilitate the scaling of the fabric for different processor designs or configurations. In some embodiments, the fabric supports different virtual channels (e.g., non - stallable channels and stallable channels) that can have different quality - of - service parameters.

[0021] Within a given tile, the fabric can support single - cycle arbitration between inputs, which can be achieved by performing various arbitration operations at least partially in parallel for different inputs. The fabric can also utilize a multi - cycle window priority update scheme based on arbitration results.

[0022] Figure 1A and Figure 1B The following discussion of [and] provides an overview of an example graphics processor, but the disclosed circuits can be used in various types of processors (e.g., central processing units, microcontrollers, machine - learning accelerators, etc.). According to some embodiments, Figures 2 to 10 the discussion of [and] covers various tiled fabric circuits and arbitration techniques.

[0023] Overview of Graphics Processing

[0024] Reference Figure 1A Note: There seems to be an unclear part in the original text where it says "and" without clear context in lines 33 and 35. I've translated it as "[and]" for now. If you can clarify the original meaning, a more accurate translation can be provided., shows a flowchart illustrating an example processing flow 100 for processing graphics data. In some embodiments, the transformation and lighting routine 110 may involve processing lighting information for vertices received from an application based on defined light source positions, reflectivity, etc., assembling the vertices into polygons (e.g., triangles), and transforming the polygons to the correct size and orientation based on their positioning in three-dimensional space. The clipping routine 115 may involve discarding polygons or vertices that are outside the visible region. In some embodiments, geometric processing may be performed using object shaders and mesh shaders prior to rasterization for flexibility and efficient processing. The rasterization routine 120 may involve defining fragments within each polygon and assigning initial color values to each fragment based on, for example, the texture coordinates of the polygon vertices. Fragments may specify the attributes of the pixels they overlap, but the actual pixel attributes may be determined based on combining multiple fragments (e.g., in a frame buffer), ignoring one or more fragments (e.g., if they are covered by other objects), or both. The shading routine 130 may involve changing pixel components based on lighting, shadows, bump mapping, translucency, etc. The shaded pixels may be assembled in the frame buffer 135. Modern GPUs typically include programmable shaders that allow application developers to customize the shading and other processing routines. Thus, in various embodiments, Figure 1A the example elements may be executed in various orders, executed in parallel, or omitted. Additional processing routines may also be implemented.

[0025] Now refer to Figure 1B , which shows a simplified block diagram of an illustrative graphics unit 150 according to some embodiments. In the illustrative embodiment, the graphics unit 150 includes a programmable shader 160, a vertex pipe 185, a fragment pipe 175, a texture processing unit (TPU) 165, an image write buffer 170, and a memory interface 180. In some embodiments, the graphics unit 150 is configured to process both vertex data and fragment data using the programmable shader 160, which may be configured to process graphics data in parallel using multiple execution pipelines or instances.

[0026] In the illustrative embodiment, the vertex pipe 185 may include various fixed-function hardware configured to process vertex data. The vertex pipe 185 may be configured to communicate with the programmable shader 160 to coordinate vertex processing. In the illustrative embodiment, the vertex pipe 185 is configured to send the processed data to the fragment pipe 175 or the programmable shader 160 for further processing.

[0027] In an illustrative embodiment, the fragment pipe 175 may include various fixed-function hardware configured to process pixel data. The fragment pipe 175 may be configured to communicate with the programmable shader 160 to coordinate fragment processing. The fragment pipe 175 may be configured to perform rasterization on polygons from the vertex pipe 185 or the programmable shader 160 to generate fragment data. The vertex pipe 185 and the fragment pipe 175 may be coupled to a memory interface 180 (coupling not shown) to access graphics data.

[0028] In an illustrative embodiment, the programmable shader 160 is configured to receive vertex data from the vertex pipe 185 and fragment data from the fragment pipe 175 and the TPU 165. The programmable shader 160 may be configured to perform vertex processing tasks on the vertex data, which may include various transformations and adjustments of the vertex data. For example, in an illustrative embodiment, the programmable shader 160 is further configured to perform fragment processing tasks on the pixel data, such as texturing and shading. The programmable shader 160 may include multiple sets of multiple execution pipelines for parallel processing of data.

[0029] In some embodiments, the programmable shader includes pipelines configured to execute one or more different SIMD groups in parallel. Each pipeline may include various stages configured to perform operations (such as fetch, decode, issue, execute, etc.) during a given clock cycle. The concept of a processor "pipeline" is well understood and refers to the concept of dividing the "work" of a processor's instruction execution into multiple stages. In some embodiments, decoding, dispatching, execution (i.e., doing), and retirement of instructions may be examples of different pipeline stages. Many different pipeline architectures may have different orderings of elements / parts. The various pipeline stages perform such steps on an instruction during one or more processor clock cycles and then pass the instruction or the operation associated with the instruction to other stages for further processing.

[0030] The term "SIMD group" is intended to be interpreted according to its well-known meaning, which includes a group of threads for which the processing hardware processes the same instructions in parallel using different input data for different threads. A SIMD group may also be referred to as a SIMT (single instruction multiple thread group), single instruction parallel thread (SIPT), or channel stack thread. Various types of computer processors may include multiple sets of pipelines configured to execute SIMD instructions. For example, a graphics processor typically includes programmable shader cores that are configured to execute instructions for a group of related threads in a SIMD fashion. Other examples of names that may be used for SIMD groups include: wavefront, clique, or warp. A SIMD group may be part of a larger group of threads, which may be split into multiple SIMD groups based on the parallel processing capabilities of the computer. In some embodiments, each thread is assigned to a hardware pipeline (which may be referred to as a "channel"), which fetches the operands of the thread and executes the specified operations in parallel with other pipelines of the group of threads. Note that a processor may have a large number of pipelines such that multiple separate SIMD groups may also execute in parallel. In some embodiments, each thread has a private operand storage, for example, in a register file. Thus, reading a particular register from the register file provides a version of the register for each thread in the SIMD group.

[0031] As used herein, the term "thread" includes its well-known meaning in the art and refers to a sequence of program instructions that can be scheduled to execute independently of other threads. A SIMD group may include multiple threads to execute in lockstep. Multiple threads may be included in a task or process, which may correspond to a computer program. The threads of a given task may or may not share resources such as registers and memory. Thus, a context switch may or may not be performed when switching between the threads of the same task.

[0032] In some embodiments, multiple programmable shader units 160 are included in the GPU. In these embodiments, the global control circuit may assign work to different sub-parts of the GPU, which in turn may allocate the work to the shader cores for processing by the shader pipelines.

[0033] In an illustrative embodiment, the TPU 165 is configured to schedule fragment processing tasks from the programmable shader 160. In some embodiments, the TPU 165 is configured to prefetch texture data and assign an initial color to the fragments for further processing by the programmable shader 160 (e.g., via the memory interface 180). The TPU 165 may be configured to provide fragment components in, for example, a normalized integer format or a floating-point format. In some embodiments, the TPU 165 is configured to provide fragments in groups of four (“fragment quads”) in a 2x2 format, which are processed by a set of four execution pipelines in the programmable shader 160.

[0034] In some embodiments, the image write buffer 170 is configured to store processed tiles of an image and may perform operations on the rendered image before transmitting it for display or sending it to memory for storage. In some embodiments, the graphics unit 150 is configured to perform tiled deferred rendering (TBDR). In tiled rendering, different portions of screen space (e.g., squares or rectangles of pixels) may be processed separately. In various embodiments, the memory interface 180 may facilitate communication with one or more of various memory hierarchies.

[0035] As discussed above, a graphics processor typically includes specialized circuitry configured to perform certain graphics processing operations requested by a computing system. For example, this may include fixed-function vertex processing circuitry, pixel processing circuitry, or texture sampling circuitry. The graphics processor may also perform non-graphics computing tasks that may use GPU shader cores but not fixed-function graphics hardware. As an example, machine learning workloads (which may include inference, training, or both) are typically assigned to the GPU due to the parallel processing capabilities of the GPU. Thus, the compute kernels executed by the GPU may include program instructions specifying machine learning tasks such as implementing a neural network layer or other aspects of a machine learning model to be executed by the GPU shader. In some cases, non-graphics workloads may also use the specialized graphics circuitry for purposes different from those originally intended.

[0036] Additionally, in other embodiments, the various circuits and techniques discussed herein with reference to graphics processors may be implemented in other types of processors. Other types of processors may include general-purpose processors such as CPUs or machine learning or artificial intelligence accelerators with dedicated parallel processing capabilities. These other types of processors may not be configured to execute graphics instructions or perform graphics operations. For example, other types of processors may not include the fixed-function hardware included in a typical GPU. A machine learning accelerator may include dedicated hardware for certain operations such as implementing neural network layers or other aspects of a machine learning model. Generally, there may be design trade-offs among memory requirements, computational power, power consumption, and programmability of a machine learning accelerator. Accordingly, different embodiments may focus on different performance goals. A developer may choose from multiple potential hardware targets for a given machine learning application, such as from a general-purpose processor, a GPU, and different dedicated machine learning accelerators.

[0037] Overview of Tiled Structures and Example Tiled Structures

[0038] Figure 2 FIG. 7 is a block diagram illustrating an example fabric circuit 210 that couples multiple clients in accordance with some embodiments. In the illustrated embodiment, the clients include a memory interface 220, a cache level A 230, a cache level B 250, a shader data path circuit 240, and a fixed-function circuit 260.

[0039] The memory interface 220 may provide access to system memory. In some embodiments, the interface 220 provides access to a unified memory subsystem that is shared and dynamically allocable. The various clients may utilize the memory space via the interface 220. Caches may use the memory interface 220 to store data in a backing memory, e.g., for eviction or write-through operations. Registers and dedicated memory spaces for the various clients may be memory backed via the memory interface 220.

[0040] Certain clients may also communicate directly with each other, e.g., for non-memory access such as control packets. Generally, the fabric circuit 210 may support various suitable packets via a packet-switched network. In a packet-switched implementation, the fabric circuit 210 may route packets from a source to a destination based on header information and may not be aware of the packet payload content.

[0041] The client may include any of various cache levels (e.g., 230 and 250) that may be dedicated to instructions / data or shared for both. For example, the various caches may be write-back or write-through implementations. The shader data path circuit 240 may include SIMD execution pipelines configured to execute various programs, including compute programs, vertex programs, fragment programs, etc. The fixed function circuits may be configured to perform various operations, such as texture sampling operations, ray tracing operations, or specific vertex operations. In other embodiments, various clients are envisioned in addition to or instead of the illustrated client.

[0042] In some embodiments, one or more clients may have access to other structures, such as a system-on-chip structure that communicates with elements external to the processor including the structure circuit 210. Thus, the structure circuit 210 may be one or more structures of a given computing system, and different structures may have different configurations.

[0043] Examples of communication via the structure include: communication with the level 0 data cache, requests for vertex data from the level 1 cache, fetch requests and fragment feedback data, initializing the SIMD group state in dedicated memory, row fill requests from caches (including data caches, instruction caches, or both), etc. Generally, some paths through the structure may be considered more important than others, and these paths may be designed to potentially provide lower latency, different power consumption, etc. to those paths at the expense of other paths.

[0044] Figure 3 is a block diagram illustrating an example tiled structure architecture according to some embodiments. In the illustrated embodiment, the structure circuit 210 includes tiles 310A - 310M. In this example, each tile includes client inputs and outputs, inputs and outputs to other tiles, internal slices 320A - 320N, and internal interfaces between the slices. As briefly discussed above, different tiles may have different numbers of client inputs and outputs (likewise, different clients may have different numbers of inputs or outputs to the structure). Different tiles in the structure may also have different numbers or configurations of internal slices.

[0045] As shown, the tiles may be arranged in a chain configuration where a given tile communicates with its left and right neighboring tiles (except for the tiles at the ends of the chain, which may have a single neighboring tile). In other embodiments, various tile connection topologies (e.g., ring, star, etc.) may be implemented. Thus, generally, a given tile may receive inputs from one or more other tiles and send outputs to one or more other tiles.

[0046] In some embodiments, the fabric performs single-cycle arbitration between requests for a given tile based on the current priority of the inputs and may adjust the priority over a multi-cycle window.

[0047] Example Tiled Circuit

[0048] Figure 4 is a block diagram illustrating a more detailed view of a given tile according to some embodiments. In the illustrated embodiment, the tile includes slices 320A - 320N, network interfaces (NIs) 430A - 430M and 440A - 440Q, and buffers 450A and 450B.

[0049] Generally, a given tile 310 has multiple inputs and multiple resources that can be assigned to those inputs. For example, the resources can include outputs in various directions and links within the tile 310.

[0050] Note that various embodiments herein are described using north / south / east / west or horizontal / vertical coordinates. These terms are used to describe example topologies but are not intended to limit the positioning of the various inputs and outputs. Additionally, topologies with other numbers of dimensions can be utilized, such as tiles with 2, 3, 5, 6, etc. sides.

[0051] As indicated by the dashed lines, slice 320 can include dedicated resources for non-stallable virtual channels and arbitration resources for stallable virtual channels. In some embodiments, all resources are arbitrated, but the highest K priority values are dedicated to non-stallable channels, and a given tile includes sufficient resources to always satisfy requests from the highest K priority values (even in these embodiments, specific resources may not be assigned to those channels and dedicated resources may not be physically separated from non-dedicated resources). As discussed in detail below after the description in Figure 4 a detailed example slice is shown in Figure 5 is shown.

[0052] In the illustrated embodiment, network interfaces 430 and 440 provide interfaces to both north and south clients. Different interfaces can have different configurations, for example to conform to different client parameters. The receive (Rx) and transmit (Tx) network interfaces can be configured similarly or differently.

[0053] In the illustrated embodiment, buffer 450 stores packet data output to other tiles in the east and west directions. As shown, in this embodiment, tile 310 also receives inputs from the east and west directions (e.g., via buffers storing output packets from other tiles).

[0054] As shown, slices can be arranged according to the chain topology within a tile, although other topologies are conceivable. In some embodiments, the tile can include crossbar circuitry (not shown) at each end of the chain to route communications between tiles (e.g., between the tile input buffer / output buffer and adjacent slice circuitry).

[0055] In some embodiments, each input to a tile is one of the following: a North Tx NI input for a packet entering the network from the north in the tile, a South Tx NI input for a packet entering the network from the south in the tile, an east input buffer for a packet arriving at the east of the tile from the tile, and a west input buffer for a packet arriving at the west of the tile from the tile.

[0056] In some embodiments, each input interface to a tile has a unit input valid signal set when a packet is available on that interface and a unit output ready signal set by the tile when the packet has been assigned through all the required routes of the tile. In some embodiments, each input interface to a tile has an input resource requirement flag that indicates the type of resources required to route the packet from the input of the current tile to the output.

[0057] As an example, the types of resources that an input packet may require can include, but are not limited to: a connection to the North Rx NI of the tile, a connection to the South Rx NI of the tile, a horizontal link through the east of the tile, an east output of the tile; a horizontal link through the west of the tile, and a west output of the tile.

[0058] Note that in some embodiments, various restrictions can be placed on the input resource requirements. For example, in some embodiments, the north input may not require the North Rx NI within the same slice, the south input may not require the South Rx NI within the same slice, the east input may not require the east horizontal link or the east output, the west input may not require the west horizontal link or the west output, each valid input must require at least one north or south Rx NI or east or west output, if the input requires an east output then it requires an east horizontal link, if the Tx NI input requires an Rx NI in the slice to the east of the input then it requires an east horizontal link, if the west input requires an Rx NI in any slice other than the westernmost input then it requires an east horizontal link, if the input requires a west output then it requires a west horizontal link, if the Tx NI input requires an Rx NI in the slice to the west of the input then it requires a west horizontal link, and if the east input requires an Rx NI in any slice other than the easternmost input then it requires a west horizontal link. The specific restrictions are included for purposes of explanation (and to provide background for the arbitration techniques discussed below) and are not intended to limit the scope of the present disclosure.

[0059] Example Slicing Circuit

[0060] Figure 5 is a block diagram illustrating a detailed example slice according to some embodiments. In the illustrated example, the slice includes multiplexers (MUXs) 510, 515, 520A - 520N, and 530A - 530N. As shown, slice 320 has J north - client inputs, K north - client outputs, M south - client inputs, and L south - client outputs. Similarly, slice 320 has N inputs and outputs in the east and west directions. Thus, in these embodiments, each slice can have the following main parameters: the number N of east / west inputs / outputs, the number J of north inputs, the number K of north outputs, the number M of south inputs, and the number L of south outputs.

[0061] In the illustrated embodiment, MUX 510 receives all east and west inputs and selects the outputs to K north - clients. Similarly, in the illustrated embodiment, MUX 515 receives all east and west inputs and selects the outputs to L north - clients.

[0062] In the illustrated embodiment, each of MUXs 520 and 530 receives one or more horizontal inputs and multiple north inputs, multiple south - client inputs, or both. Each of MUXs 520 and 530 may be configurable by the amount of slice inputs accessible to any given horizontal output. In some cases, each horizontal MUX 520 or 530 may have at least as many inputs as the number of north inputs and south inputs plus the number of horizontal links directly in front of it. However, to improve the efficiency of horizontal - link utilization, each horizontal multiplexer may also accept inputs from up to N other horizontal links. In the case of a complete set of inputs, each slice becomes a full cross - switch, but in other embodiments (or in some slices), only a subgroup of horizontal inputs may be designated to connect to a particular MUX instance.

[0063] The control circuit can control the various illustrated multiplexers to select the input that wins arbitration to use the tile resources in a given cycle, for example, based on the arbitration techniques discussed in detail below.

[0064] Example Arbitration for Tiled Resources

[0065] Figure 6 is a flowchart illustrating an example arbitration technique 600 according to some embodiments. In some embodiments, the control circuit for each tile executes this arbitration technique in each cycle to assign resources to the inputs for that cycle. For example, the control circuit can assert control signals to the Figure 5 multiplexers based on this technique. As discussed below with reference to Figure 9 , the multi - cycle priority update procedure updates the priorities for the next set of cycles based on which inputs win arbitration within a window.

[0066] Note that Figure 7 and Figure 8 shows an example matrix that can be utilized by technique 600 and will thus be briefly referred to below in connection with Figure 6 the discussion of

[0067] At 610, in the illustrated embodiment, the control circuit determines an input priority (e.g., a unique priority identifier from 0 to N-1 for N inputs). As discussed above, a subgroup of the highest priority inputs may be reserved for one or more non-stallable channels. [[ID=X]] [[ID=X]]

[0068] At 620, in the illustrated embodiment, the control circuit forms an N-by-N priority comparison matrix (PCM), where PCM[i][j] indicates whether input-i or input-j has a higher priority (e.g., a bit may be set in the entry of the matrix if the priority of input-i is greater than or equal to the priority of input-j). For example, foreach(i,j): PCM[i][j] = i.priorityID >= j.priorityID. Note that the same PCM values may be used for multiple cycles, e.g., when the priorities have not changed between cycles.

[0069] Briefly referring to Figure 7 , an example PCM 700 is shown, where input A has a priority of 5, input B has a priority of 7, and input N-1 has a priority of 2. Note that only a portion of the PCM may actually be stored in the circuit, e.g., because the diagonal may be irrelevant (x) and the two other portions may be mirrored along the diagonal. In this example, PCM[1][0] is set because input B has a higher priority than input A.

[0070] At 630, in the illustrated embodiment, the control circuit forms an N-by-N priority mask matrix (PMM) for each tile resource R. In this example, if both inputs request resource R and input-i has a priority greater than or equal to input-j, then PMM[i][j] is set. Specifically, foreach(r): r.PMM[i][j] = i.valid && i.requiresResource[r] && j.valid && j.requiresResource[r] && PCM[i][j]. Note that, for example, the control circuit may determine which inputs target which resources based on packet header information indicating the destination client circuit.

[0071] Briefly referring to Figure 8 , an example PMM 800 is shown. In this example, the priorities are the same as for Figure 7have the same priority, and Input A and N request resource R (while Input B does not). In this example, PMM[0][N-1] is set because both Input A and Input N-1 request resource R and Input A has a higher priority. Note that this data structure can be stored in a compressed form similar to PCM700, as discussed above, for example, because the diagonal values are known and the remainder can be mirrored.

[0072] At 640, in the illustrated embodiment, the control circuit sums across each row of the PMM to generate a sub-resource allocation ID for each input for each resource. For example, foreach(i,j): i.allocationID += r.PMM[i][j].

[0073] Referring again to Figure 8 , the sub-resource allocation ID for Input A is the sum of the row of Input A (the top row). In this embodiment, a sum of zero indicates that Input A has not requested the resource. A sum of 1 means that A is the highest priority input requesting the resource. A sum of N means that A is the lowest priority input requesting the resource (all other inputs have higher priority and their bits are in the row group).

[0074] At 650, in the illustrated embodiment, the control circuit performs a resource availability count to determine the number of sub-resources available per resource. This can be based on buffer status, the number of outputs of one or more types of tiles, etc.

[0075] At 660, in the illustrated embodiment, the control circuit compares the sub-resource ID for each input for each resource with the availability count for that resource. For example, in parallel, foreach(r,s): r.availabilityCount += r.subResourceAvailable[s]. If the ID does not exceed the count, for example, foreach(i,r): i.allocationLegal[r] = i.allocationID[r] <= r.availabilityCount, then the control circuit determines a legal allocation. The control circuit can select inputs to provide data for a legal allocation, or otherwise stall the inputs.

[0076] In this way, the assignment of inputs to resources can change each cycle, even for cycles with the same input priorities, for example because the inputs that actually request the resource can vary in different cycles. Additionally, as discussed above, various computations (such as determining the PMM for each resource and determining the resource availability count) can be performed in parallel to facilitate single-cycle arbitration (or more generally, arbitration using a small number of cycles).

[0077] Example Priority Update Procedure

[0078] Figure 9 It is a flowchart illustrating an example priority update technique 900 according to some embodiments. In some embodiments, a priority update window is defined as, for example, a specific number of cycles, and the control circuit updates the priority of the inputs for a given tile based on the arbitration decisions during the window.

[0079] At 910, in the illustrated example, priority IDs 0 through K-1 are reserved for K non-stallable inputs. These priority IDs can be dynamically assigned to each priority update procedure or can be fixed. In other embodiments, a field can indicate whether an input is non-stallable and non-stallable inputs can be considered to have a higher priority than any stallable input.

[0080] At 920, in the illustrated example, the control circuit determines, over M cycles, the winner, loser, and non-requester among the stallable inputs. For example, the winner can be the input whose requests are fully allocated the requested resources. In some embodiments, the winner can be divided into a full winner (the input whose most recent valid request is fully allocated the requested resources) and a partial winner (an input having a previous request that was fully allocated all requested resources within the decision window, but whose most recent valid request was not fully allocated all requested resources within the window). The loser can be the input for which no valid request was fully allocated the requested resources within the decision window. The invalid / non-requester can be the input for which no valid request was received within the decision window.

[0081] At 930, in the illustrated example, the control circuit adjusts the priorities such that the highest priority winner becomes the lowest priority stallable input and the highest priority loser becomes the highest priority stallable input. The invalid inputs can have a resulting priority value between the loser and the winner. In embodiments that track partial winners, the partial winners can have a new priority value between the invalid inputs and the winner. Additionally, the control circuit can reverse the priorities within a given class of inputs, e.g., reverse the priority order for winner inputs and partial winner inputs.

[0082] Generally speaking, the disclosed techniques can advantageously provide fairness over time among stallable inputs. Additionally, in some embodiments, the priority can only be decreased for non-stallable inputs that have been granted at least one request within the window, thus advantageously ensuring forward progress.

[0083] Example Method

[0084] Figure 10 It is a flowchart illustrating an example method for using a tiled structure according to some embodiments. Figure 10The methods shown herein can be used in conjunction with any of the computer circuits, systems, devices, components, or assemblies disclosed herein. In various embodiments, some of the method elements shown may be executed concurrently in a different order than shown, or may be omitted. Additional method elements may also be executed as needed.

[0085] At 1010, in the illustrated embodiment, the computing system assigns communication resources of the tiling structure circuit for a given cycle based on priority information of the inputs for a given tile instance.

[0086] In some embodiments, the structure includes at least a first instance and a second instance of a tile, and the tile includes: a client input configured to interface with a client circuit of the computing system, a tile input configured to interface with one or more other tile instances, and communication resources that can be assigned to the client input and the tile input. The communication resources may include: a plurality of internal links, a client output configured to interface with the client circuit, and a tile output configured to interface with one or more other tile instances.

[0087] In some embodiments, the assignment assigns the communication resources of the tile instance to at least a portion of the client input and the tile input for a next cycle based on priority information of the inputs of the given tile instance.

[0088] In some embodiments, to assign the communication resources of a given tile instance, the control circuit is configured to: determine, at least partially in parallel, priority information of the inputs for the tile instance for a plurality of communication resources and a plurality of inputs, determine the number of inputs that have a higher priority than a given input and request the same resource, determine the number of inputs that the resource can service in a given cycle for a given resource, and assign the communication resources to the inputs for the given cycle based on the determination of the number of inputs.

[0089] In some embodiments, the communication resources include dedicated resources for one or more non-stallable virtual channels and arbitration resources for stallable virtual channels. In some embodiments, the control circuit is configured to assign a set of highest priority indications to non-stallable virtual channel inputs.

[0090] At 1020, in the illustrated embodiment, the computing system updates the priority information of a given tile instance of the structure circuit based on the assignment result over a plurality of cycles.

[0091] In some embodiments, to update the priority information for a given tile instance based on the assignment results over multiple cycles, the control circuit classifies the inputs to the tile instance based on whether the requested resources are received in whole or in part over multiple cycles, and updates the priority of the inputs to the tile instance based on the classification. For example, the categories may include: a winner input category, where the most recent valid request of the winner input category is fully allocated all requested resources; a partial winner input category, where the most recent valid request of the partial winner input category is not fully allocated all requested resources over multiple cycles, and for which a previous request was fully allocated all requested resources over the multiple cycles; a loser category, for which no valid request is fully allocated the requested resources over the multiple cycles; and an invalid input category, for which no valid request is received over the multiple cycles. In some embodiments, to update the priorities, the control circuit performs the prioritization in the following order from highest to lowest priority: loser inputs, invalid inputs, partial winner inputs, and then winner inputs. In some embodiments, the control circuit may reverse the prioritization among partial winner inputs and reverse the prioritization among winner inputs.

[0092] In some embodiments, a given tile instance includes a plurality of slices, and the plurality of internal links are links between the slices. In some embodiments, the slices are arranged in a chain topology, and the given tile instance includes cross switches at one or both ends of the tile chain (e.g., included as part of the end slices such that the multiplexers of those slices have additional inputs for forming a full cross switch). In some embodiments, the cross switches are connected to buffers configured to store data for the tile outputs and data from the tile inputs.

[0093] In some embodiments, a first instance and a second instance of a tile include a different number of inputs and a different amount of communication resources. In some embodiments, the fabric circuit includes a chain of tile instances that includes the first instance and the second instance of the tile, where adjacent tile instances in the chain are connected via at least a portion of the tile inputs and tile outputs of a given tile instance.

[0094] As used herein, the terms "clock" and "clock signal" refer to a periodic signal, e.g., as in a binary (two-valued) electrical signal. The clock changes periodically between "levels" of the clock, such as the voltage range of an electrical signal. For example, a voltage greater than 0.7 volts may be used to represent one clock level, and a voltage less than 0.3 volts may be used to represent another level in a binary configuration. As used herein, the term "clock edge" refers to a change in the clock signal from one level to another. As used herein, in the context of a clock signal, the term "switch" refers to changing the value of the clock signal from one level to another in a binary clock configuration. As used herein, the term clock "pulse" refers to the interval of the clock signal between successive edges of the clock signal (e.g., the interval between a rising edge and a falling edge or between a falling edge and a rising edge). As used herein, the term "clock period" refers to the interval between the edges of the clock (e.g., in a conventional implementation, between rising edges or between falling edges, or in a dual-edge implementation, between adjacent edges).

[0095] Example Device

[0096] Now referring to Figure 11 , a block diagram illustrating an exemplary embodiment of an exemplary device 1100 is shown. In some embodiments, the elements of device 1100 may be included within a system-on-chip. In some embodiments, device 1100 may be included in a mobile device that may be battery-powered. Thus, power consumption of device 1100 may be an important design consideration. In the illustrated embodiment, device 1100 includes fabric 1110, a compute complex 1120, an input / output (I / O) bridge 1150, a cache / memory controller 1145, a graphics unit 1175, and a display unit 1165. In some embodiments, in addition to or instead of the components shown, device 1100 may include other components (not shown), such as video processor encoders and decoders, image processing or recognition elements, computer vision elements, and the like.

[0097] Fabric 1110 may include various interconnects, buses, MUXes, controllers, etc., and may be configured to facilitate communication between the various elements of device 1100. In some embodiments, portions of fabric 1110 may be configured to implement various different communication protocols. In other embodiments, fabric 1110 may implement a single communication protocol, and elements coupled to fabric 1110 may internally translate from the single communication protocol to other communication protocols. Note that fabric 1110 may have a different topology and configuration than fabric circuit 210.

[0098] In the illustrated embodiment, the compute complex 1120 includes a bus interface unit (BIU) 1125, caches 1130, and cores 1135 and 1140. In various embodiments, the compute complex 1120 may include various numbers of processors, processor cores, and caches. For example, the compute complex 1120 may include 1, 2, or 4 processor cores, or any other suitable number. In one embodiment, the cache 1130 is a set of associative L2 caches. In some embodiments, the cores 1135 and 1140 may include internal instruction and data caches. In some embodiments, a coherence unit (not shown) in the fabric 1110, the cache 1130, or elsewhere in the device 1100 may be configured to maintain coherence among the various caches of the device 1100. The BIU 1125 may be configured to manage communication between the compute complex 1120 and other elements of the device 1100. Processor cores such as cores 1135 and 1140 may be configured to execute instructions of a particular instruction set architecture (ISA) that may include operating system instructions and user application instructions.

[0099] The cache / memory controller 1145 may be configured to manage data transfer between the fabric 1110 and one or more caches and memories. For example, the cache / memory controller 1145 may be coupled to an L3 cache, which in turn may be coupled to system memory. In other embodiments, the cache / memory controller 1145 may be directly coupled to memory. In some embodiments, the cache / memory controller 1145 may include one or more internal caches.

[0100] As used herein, the term "coupled to" may indicate one or more connections between elements, and coupling may include intermediate elements. For example, in Figure 11 FIG., the graphics unit 1175 may be described as "coupled to" memory through the fabric 1110 and the cache / memory controller 1145. In contrast, in the Figure 11 illustrated embodiment of FIG., the graphics unit 1175 is "directly coupled" to the fabric 1110 because there are no intermediate elements.

[0101] The graphics unit 1175 may include one or more processors, e.g., one or more graphics processing units (GPUs). For example, the graphics unit 1175 may receive graphics-oriented instructions such as Metal or Instructions. The graphics unit 1175 may execute specialized GPU instructions or perform other operations based on the received graphics-oriented instructions. The graphics unit 1175 may generally be configured to process large chunks of data in parallel and may build an image in a frame buffer for output to a display, which may be included in the device or may be a separate device. The graphics unit 1175 may include transform, lighting, triangle, and rendering engines in one or more graphics processing pipelines. The graphics unit 1175 may output pixel information for displaying an image. In various embodiments, the graphics unit 1175 may include programmable shader circuitry that may include highly parallel execution cores configured to execute graphics programs, which may include pixel tasks, vertex tasks, and compute tasks (which may or may not be graphics-related).

[0102] In some embodiments, the disclosed communication fabric may be included in the graphics unit 1175 and may be used for communication between data path circuitry, one or more cache levels, fixed function circuitry, etc. The disclosed techniques may advantageously facilitate communication in a scalable manner. Note that the disclosed fabric may also or alternatively be implemented in other processors such as the compute complex 1120.

[0103] The display unit 1165 may be configured to read data from the frame buffer and provide a stream of pixel values for display. In some embodiments, the display unit 1165 may be configured as a display pipeline. Additionally, the display unit 1165 may be configured to blend multiple frames to produce an output frame. Additionally, the display unit 1165 may include one or more interfaces (e.g., or an embedded display port (eDP)) for coupling to a user display (e.g., a touchscreen or an external display).

[0104] The I / O bridge 1150 may include various components configured to implement, for example, universal serial bus (USB) communication, security, audio, and low-power always-on functionality. The I / O bridge 1150 may also include interfaces such as pulse width modulation (PWM), general-purpose input / output (GPIO), serial peripheral interface (SPI), and inter-integrated circuit (I2C). Various types of peripheral devices and devices may be coupled to the device 1100 via the I / O bridge 1150.

[0105] In some embodiments, device 1100 includes a network interface circuit (not explicitly shown) that may be connected to fabric 1110 or I / O bridge 1150. The network interface circuit may be configured to communicate via various networks, which may be wired networks, wireless networks, or both. For example, the network interface circuit may be configured to communicate via a wired local area network, a wireless local area network (e.g., via WiFi), or a wide area network (e.g., the Internet or a virtual private network). In some embodiments, the network interface circuit is configured to communicate via one or more cellular networks using one or more radio access technologies. In some embodiments, the network interface circuit is configured to communicate using device-to-device communication (e.g., Bluetooth or WiFi Direct), etc. In various embodiments, the network interface circuit provides device 1100 with connections to various types of other devices and networks.

[0106] Example Application

[0107] Turning now to Figure 12 , various types of systems are shown that may include any of the circuits, devices, or systems described above. Systems or devices 1200 that may incorporate or otherwise utilize one or more of the techniques described herein can be used in a wide range of fields. For example, systems or devices 1200 can be used as part of the hardware of systems such as desktop computer 1210, laptop computer 1220, tablet computer 1230, cellular or mobile phone 1240, or television 1250 (or a set-top box coupled to a television).

[0108] Similarly, the disclosed elements can be used in wearable device 1260, such as a smartwatch or a health monitoring device. In many embodiments, a smartwatch can implement a variety of different functions—for example, access to email, cellular service, calendar, health monitoring, etc. Wearable devices can also be designed to perform only health monitoring functions, such as monitoring a user's vital signs, performing epidemiological functions such as contact tracing, providing communication to emergency medical services, etc. Other types of devices are also envisioned, including devices worn around the neck, devices implantable in the human body, glasses or helmets designed to provide computer-generated reality experiences, such as those based on augmented reality and / or virtual reality, etc.

[0109] System or device 1200 can also be used in a variety of other environments. For example, system or device 1200 can be used in the context of a server computer system (such as a dedicated server) or on shared hardware that implements cloud-based service 1270. Further, system or device 1200 can be implemented in a wide range of dedicated everyday devices, including devices 1280 common in the home, such as refrigerators, thermostats, security cameras, etc. The interconnection of such devices is commonly referred to as the "Internet of Things" (IoT). The elements can also be implemented in various modes of transportation. For example, system or device 1200 can be used for control systems, guidance systems, entertainment systems, etc. of various types of vehicles 1290.

[0110] Figure 12 The applications shown are merely exemplary and are not intended to limit the potential future applications of the disclosed system or device. Other example applications include, but are not limited to: portable gaming devices, music players, data storage devices, unmanned aerial vehicles, etc.

[0111] Example Computer - Readable Medium

[0112] The present disclosure has described various example circuits in detail above. It is intended that the present disclosure cover not only embodiments including such circuits, but also computer-readable storage media including design information specifying such circuits. Thus, the present disclosure is intended to support claims covering not only devices including the disclosed circuits, but also storage media that specify the circuits in a format recognized by a manufacturing system configured to produce hardware (e.g., integrated circuits) including the disclosed circuits. Claims to such storage media are intended to cover, for example, entities that generate circuit designs but do not themselves manufacture the design.

[0113] Figure 13 is a block diagram showing an example non-transitory computer-readable storage medium storing circuit design information. In the illustrated embodiment, semiconductor manufacturing system 1320 is configured to process design information 1315 stored on non-transitory computer-readable medium 1310 and manufacture integrated circuit 1330 based on design information 1315.

[0114] The non-transitory computer-readable storage medium 1310 may include any one of various suitable types of memory devices or storage devices. The non-transitory computer-readable storage medium 1310 may be an installation medium, such as a CD-ROM, a floppy disk, or a magnetic tape device; a computer system memory or random access memory, such as DRAM, DDR RAM, SRAM, EDO RAM, Rambus RAM, etc.; a non-volatile memory, such as a flash memory, a magnetic medium, e.g., a hard disk drive or an optical storage device; a register, or other similar types of memory elements, etc. The non-transitory computer-readable storage medium 1310 may also include other types of non-transitory memories or combinations thereof. The non-transitory computer-readable storage medium 1310 may include two or more memory media that may reside at different locations, e.g., in different computer systems connected by a network.

[0115] The design information 1315 may be specified in any of various suitable computer languages, including hardware description languages such as, but not limited to: VHDL, Verilog, SystemC, SystemVerilog, RHDL, M, MyHDL, etc. The design information 1315 may be capable of being used by the semiconductor manufacturing system 1320 to fabricate at least a portion of the integrated circuit 1330. The format of the design information 1315 may be recognized by at least one semiconductor manufacturing system 1320. In some embodiments, the design information 1315 may also include one or more cell libraries that specify the synthesis, layout, or both of the integrated circuit 1330. In some embodiments, the design information is specified, in whole or in part, in the form of a netlist that specifies cell library elements and their connectivity. The separately obtained design information 1315 may or may not include sufficient information for fabricating the corresponding integrated circuit. For example, the design information 1315 may specify the circuit elements to be fabricated, but not their physical layout. In such a case, the design information 1315 may need to be combined with layout information to actually fabricate the specified circuit.

[0116] In various embodiments, the integrated circuit 1330 may include one or more custom macro cells, such as memories, analog or mixed-signal circuits, etc. In such a case, the design information 1315 may include information related to the included macro cells. Such information may include, but is not limited to, a schematic capture database, mask design data, behavioral models, and device or transistor-level netlists. As used herein, the mask design data may be formatted according to the Graphics Data System (GDSII) or any other suitable format.

[0117] The semiconductor manufacturing system 1320 may include any of a variety of suitable elements configured to manufacture integrated circuits. This may include, for example, elements for depositing semiconductor materials (e.g., on a wafer that may include a mask), removing materials, shaping the deposited materials, modifying the materials (e.g., by doping the materials or using ultraviolet treatment to modify the dielectric constant), etc. The semiconductor manufacturing system 1320 may also be configured to perform various tests on the manufactured circuits for proper operation.

[0118] In various embodiments, the integrated circuit 1330 is configured to operate according to a circuit design specified by the design information 1315, which may include performing any of the functions described herein. For example, the integrated circuit 1330 may include Figure 1B , Figures 2 to 5 and Figure 11 any of the various elements shown. Additionally, the integrated circuit 1330 may be configured to perform various functions described herein in connection with other components. Additionally, the functionality described herein may be performed by multiple connected integrated circuits.

[0119] As used herein, a phrase of the form "design information specifying a design of a circuit configured to..." does not imply that the circuit in question must be manufactured in order to meet the element. Instead, the phrase indicates that the design information describes a circuit that, when manufactured, will be configured to perform the indicated actions or will include the specified components.

[0120] ***

[0121] This disclosure includes references to "an embodiment" or a group of "embodiments" (e.g., "some embodiments" or "various embodiments"). An embodiment is a different specific implementation or instance of the disclosed concept. References to "an embodiment", "one embodiment", "a particular embodiment", etc. do not necessarily refer to the same embodiment. A large number of possible embodiments are envisioned, including those specifically disclosed, as well as modifications or alternatives that fall within the spirit or scope of this disclosure.

[0122] The present disclosure may discuss potential advantages that may result from the disclosed embodiments. Not all implementations of these embodiments will necessarily exhibit any or all of the potential advantages. Whether a particular implementation realizes an advantage depends on many factors, some of which are outside the scope of the present disclosure. In fact, there are many reasons why a particular implementation falling within the scope of the claims may not exhibit some or all of the disclosed advantages. For example, a particular implementation may include other circuitry outside the scope of the present disclosure that, when combined with one of the disclosed embodiments, negates or diminishes one or more of the disclosed advantages. Additionally, suboptimal design implementation of a particular implementation (e.g., a particular implementation technique or tool) may also negate or diminish the disclosed advantages. Even assuming an implementation of the technology, the realization of an advantage may still depend on other factors, such as the circumstances of the environment in which the implementation is deployed. For example, the input provided to a particular implementation may prevent one or more of the problems addressed in the present disclosure from occurring on a particular occasion, and as a result, the benefits of its solution may not be realized. Given the existence of possible factors outside the scope of the present disclosure, it is hereby expressly stated that any potential advantages described herein should not be construed as claim limitations that must be met in order to establish infringement. Instead, the identification of such potential advantages is intended to illustrate the types of improvements available to designers who benefit from the present disclosure. Permanently describing such advantages (e.g., stating that a particular advantage “may occur”) is not intended to convey doubt as to whether such advantages can actually be realized, but rather to recognize the technical reality that the realization of such advantages generally depends on additional factors.

[0123] Unless otherwise specified, the embodiments are non - restrictive. That is, the disclosed embodiments are not intended to limit the scope of the claims drafted based on the present disclosure, even when only a single example is described with respect to a particular feature. The embodiments disclosed in the present invention are intended to be illustrative rather than restrictive, without any contrary statement in the present disclosure. Accordingly, this application is intended to allow claims that cover the disclosed embodiments, as well as such alternatives, modifications, and equivalents, which will be apparent to those skilled in the art who are aware of the beneficial effects of the present disclosure.

[0124] For example, the features in this application may be combined in any suitable manner. Accordingly, new claims may be made during the prosecution of this application (or an application claiming priority therefrom) for any such combination of features. Specifically, with reference to the appended claims, the features of dependent claims may, where appropriate, be combined with the features of other dependent claims, including claims that depend from other independent claims. Similarly, the features from corresponding independent claims may be combined, where appropriate.

[0125] Accordingly, while the appended dependent claims may be drafted such that each dependent claim depends from a single other claim, additional dependencies are also contemplated. Any combination of dependent features consistent with the present disclosure is contemplated, and such combinations may be claimed in this application or another application. In short, the combinations are not limited to those specifically recited in the appended claims.

[0126] In appropriate cases, it is also contemplated that claims drafted in one format or statutory type (e.g., apparatus) are intended to support corresponding claims in another format or statutory type (e.g., method).

[0127] ***

[0128] Because the present disclosure is a legal document, various terms and phrases may be subject to administrative and judicial interpretation. Notice is hereby given that the following paragraphs, as well as the definitions provided throughout the present disclosure, will be used to determine how claims drafted based on the present disclosure are to be interpreted.

[0129] References to items in the singular form (i.e., a noun or noun phrase preceded by "a," "an," or "the") are intended to mean "one or more" unless the context clearly dictates otherwise. Thus, without accompanying context, a reference to "an item" in a claim does not exclude additional instances of that item. "Multiple" items means a group of two or more items.

[0130] The word "may" is used herein in the permissive sense (i.e., having the potential to be able to), rather than in the mandatory sense (i.e., must).

[0131] The terms "comprising" and "including" and their forms are open-ended and mean "including but not limited to."

[0132] When the term "or" is used in the present disclosure with respect to a list of options, it will generally be understood to be used in the inclusive sense unless the context provides otherwise. Thus, the statement "x or y" is equivalent to "x or y, or both," and thus encompasses 1) x but not y, 2) y but not x, and 3) both x and y. On the other hand, phrases such as "either x or y, but not both" make it clear that "or" is used in the exclusive sense.

[0133] The phrase "w, x, y, or z, or any combination thereof" or "at least one of...w, x, y, and z" is intended to cover all possibilities of individual elements up to the total number of elements in the set. For example, given the set [w, x, y, z], these phrases cover any single element in the set (e.g., w but not x, y, or z), any two elements (e.g., w and x but not y or z), any three elements (e.g., w, x, and y but not z), and all four elements. The phrase "at least one of...w, x, y, and z" thus refers to at least one element in the set [w, x, y, z], thereby covering all possible combinations in the list of elements. This phrase should not be construed as requiring the existence of at least one instance of w, at least one instance of x, at least one instance of y, and at least one instance of z.

[0134] In the present disclosure, various "labels" may precede a noun or noun phrase. Unless the context provides otherwise, different labels for a feature (e.g., "first circuit", "second circuit", "specific circuit", "given circuit", etc.) refer to different instances of the feature. Additionally, unless otherwise specified, the labels "first", "second", and "third" do not imply any type of ordering (e.g., spatial, temporal, logical, etc.) when applied to a feature.

[0135] The phrase "based on" or "used to describe one or more factors that affect a determination. This term does not exclude the possibility that additional factors may affect the determination. That is, the determination may be based solely on the specified factors or on the specified factors and other unspecified factors. Consider the phrase "determine A based on B". This phrase specifies that B is a factor used to determine A or that B affects the determination of A. This phrase does not exclude the possibility that the determination of A may also be based on some other factor such as C. This phrase is also intended to cover embodiments where A is determined based solely on B. As used herein, the phrase "based on" is synonymous with the phrase "at least partially based on".

[0136] The phrases "responsive to" and "in response to" describe one or more factors that trigger an effect. This phrase does not exclude the possibility that additional factors may affect or otherwise trigger the effect, either in conjunction with the specified factors or independently of the specified factors. That is, the effect may be responsive solely to these factors, or it may be responsive to the specified factors and other unspecified factors. Consider the phrase "perform A in response to B". This phrase specifies that B is a factor that triggers the performance of A or triggers a particular result of A. This phrase does not exclude the possibility that the performance of A may also be responsive to some other factor, such as C. This phrase also does not exclude the possibility that the performance of A may be performed in response to B and C in conjunction. This phrase is also intended to cover embodiments where A is performed solely in response to B. As used herein, the phrase "in response to" is synonymous with the phrase "at least partially in response to". Similarly, the phrase "responsive to" is synonymous with the phrase "at least partially responsive to".

[0137] ***

[0138] Within this disclosure, different entities (which may be variously referred to as "units", "circuits", other components, etc.) may be described or claimed as "configured to" perform one or more tasks or operations. This expression - [entity] [configured to [perform one or more tasks]] - is used herein to refer to a structure (i.e., a physical thing). More specifically, this expression is used to indicate that this structure is arranged to perform one or more tasks during operation. A structure may be considered "configured to" perform a certain task even if the structure is not currently being operated. Thus, an entity described or stated as "configured to" perform a certain task refers to a physical thing for implementing that task, such as a device, a circuit, a system having a processor unit, and a memory storing executable program instructions, etc. This phrase is not used herein to refer to intangible things.

[0139] In some cases, various units / circuits / components may be described herein as performing a set of tasks or operations. It should be understood that these entities are "configured to" perform those tasks / operations even if not specifically stated.

[0140] The term "configured to" is not intended to mean "configurable to". For example, an unprogrammed FPGA is not considered to be "configured to" perform a specific function. However, the unprogrammed FPGA can be "configurable to" perform that function. After appropriate programming, the FPGA can then be considered "configured to" perform a specific function.

[0141] For the purposes of a U.S. patent application based on this disclosure, stating in a claim that a structure is "configured to" perform one or more tasks is specifically intended not to invoke 35 U.S.C. § 112(f) for that claim element. If an applicant wishes to invoke section 112(f) during the prosecution of a U.S. patent application based on this disclosure, it will use the "means for [performing a function]" structure to phrase the claim element.

[0142] Different "circuits" may be described in this disclosure. These circuits constitute hardware, which includes various types of circuit elements, such as combinational logic, clock storage devices (e.g., flip - flops, registers, latches, etc.), finite state machines, memories (e.g., random access memories, embedded dynamic random access memories), programmable logic arrays, etc. The circuits may be custom - designed or taken from a standard library. In various specific implementations, the circuits may include digital components, analog components, or a combination of both, as appropriate. Certain types of circuits may generally be referred to as "units" (e.g., decoding units, arithmetic logic units (ALUs), functional units, memory management units (MMUs), etc.). Such units also refer to circuits or circuitry.

[0143] Accordingly, the disclosed circuits / cells / components and other elements illustrated in the figures and described herein include hardware elements such as those described in the previous paragraphs. In many cases, the internal arrangement of the hardware elements in a particular circuit can be specified by describing the function of the circuit. For example, a particular "decoding unit" may be described as performing the function of "processing the opcode of an instruction and routing the instruction to one or more of a plurality of functional units", which means that the decoding unit "is configured to" perform that function. For those skilled in the art of computers, this functional specification is sufficient to imply a set of possible structures for the circuit.

[0144] In various embodiments, as described in the previous paragraphs, circuits, cells, and other elements may be defined by the functions or operations they are configured to perform. The arrangement relative to each other and such circuits / cells / components and the ways in which they interact form the microarchitecture definition of the hardware, which is ultimately fabricated in an integrated circuit or programmed into an FPGA to form the physical implementation of the microarchitecture definition. Thus, the microarchitecture definition is considered by those skilled in the art to be a structure from which many physical implementations can be derived, all of which fall within the broader structure described by the microarchitecture definition. That is, a person having a microarchitecture definition provided according to the present disclosure can, without undue experimentation and using the applications of an ordinary skilled person, implement the structure by encoding a description of the circuits / cells / components in a hardware description language (HDL) such as Verilog or VHDL. HDL descriptions are often expressed in a way that can appear functional. However, for those skilled in the art, the HDL description is a way to translate the structure of a circuit, cell, or component into the next level of implementation details. Such HDL descriptions can take the form of behavioral code (which is typically non-synthesizable), register transfer language (RTL) code (which is typically synthesizable compared to behavioral code), or structural code (e.g., a netlist specifying logic gates and their connectivity). The HDL descriptions can be sequentially synthesized for a cell library designed for a given integrated circuit manufacturing technology and can be modified for timing, power, and other reasons to obtain the final design database that is sent to the factory to generate masks and ultimately produce the integrated circuit. Some hardware circuits or portions thereof may also be custom designed in a schematic editor and captured into the integrated circuit design along with the synthesized circuits. The integrated circuit may include transistors and other circuit elements (e.g., passive elements such as capacitors, resistors, inductors, etc.), as well as interconnects between the transistors and circuit elements. Some embodiments may implement multiple integrated circuits coupled together to implement the hardware circuit, and / or discrete elements may be used in some embodiments. Alternatively, the HDL design can be synthesized into a programmable logic array such as a field programmable gate array (FPGA) and implemented in the FPGA. This decoupling between the design of a set of circuits and the subsequent low-level implementation of those circuits typically results in a situation where the circuit or logic designer never specifies a particular set of structures for the low-level implementation beyond a description of what the circuit is configured to do, because the process is performed at different stages of the circuit implementation process.

[0145] The fact that many different low-level combinations of circuit elements can be used to implement the same specification of a circuit results in a large number of equivalent structures for that circuit. As noted, these low-level circuit implementations can vary depending on manufacturing technology, the foundry selected to fabricate the integrated circuit, the cell library provided for a particular project, etc. In many cases, the choice of generating these different implementations through different design tools or methods can be arbitrary.

[0146] In addition, for a given implementation, a single implementation of a particular functional specification of a circuit typically includes a large number of devices (e.g., millions of transistors). Thus, the sheer volume of this information makes it impractical to provide a complete narrative of the low-level structures used to implement a single implementation, let alone the large number of equivalent possible implementations. For this reason, the present disclosure describes the structure of a circuit using functional shorthand commonly used in the industry.

Claims

1. An apparatus, the apparatus comprising a processor, the processor comprising: A plurality of client circuits; A fabric circuit, the fabric circuit including at least a first instance and a second instance of a tile, wherein the tile includes: A client input configured to interface with the client circuits; a tile input configured to interface with one or more other tile instances; and A communication resource that can be assigned to the client input and the tile input, wherein the communication resource includes: A plurality of internal links; A client output configured to interface with the client circuits; and A tile output configured to interface with one or more other tile instances; A control circuit configured to: In a given cycle, assign the communication resource of the tile instance to at least a portion of the client input and the tile input for the next cycle based on priority information of the input for the given tile instance; and Update the priority information for the given tile instance of the fabric circuit based on the assignment result over a plurality of cycles.

2. The apparatus according to claim 1, wherein in order to assign the communication resource of a given tile instance, the control circuit is configured to: Determine the priority information of the input for the tile instance; Determine, at least partially in parallel for a plurality of communication resources and a plurality of inputs, the number of inputs having a higher priority than a given input and requesting the same resource; Determine, for a given resource, the number of inputs that the resource can serve in the given cycle; And Assign the communication resource to the input for the given cycle based on the determination of the number of inputs.

3. The apparatus according to claim 1, wherein the communication resource includes: Dedicated resources for one or more non-stallable virtual channels; And Arbitration resources for stallable virtual channels.

4. The apparatus according to claim 3, wherein the control circuit is configured to assign a set of highest priority indications to non-stallable virtual channel inputs.

5. The apparatus according to claim 1, wherein in order to update the priority information for a given tile instance based on the assignment result over a plurality of cycles, the control circuit is configured to: Classify the inputs to the tile instance based on whether all or a portion of the requested resources are received by the inputs to the tile instance over the plurality of cycles; and Update the priority for the inputs to the tile instance based on the classification.

6. The apparatus according to claim 5, Wherein the categories of inputs include: A winner input category, the most recent valid request of which is fully allocated all requested resources; A partial winner input category, the most recent valid request of which is not fully allocated all requested resources within the plurality of cycles, and for which previous requests were fully allocated all requested resources within the plurality of cycles; A loser category, for which no valid request is fully allocated the requested resources within the plurality of cycles; And Invalid input category, for which no valid requests are received during the plurality of cycles.

7. The apparatus according to claim 6, wherein to update the priority, the control circuit is configured to prioritize in the following order from highest to lowest priority: Loser input; Invalid input; Partial winner input; and then Winner input.

8. The apparatus according to claim 7, wherein to update the priority, the control circuit is further configured to reverse the prioritization among the partial winner inputs and reverse the prioritization among the winner inputs.

9. The apparatus according to claim 1, wherein the first tile instance includes a plurality of slices, and the plurality of internal links are links between the slices.

10. The apparatus according to claim 9, wherein the slices are arranged in a chain topology, and wherein the first tile instance includes crossbar switches at one or both ends of the slice chain, wherein the crossbar switches are connected to buffers configured to store: Data for the tile output; and Data from the tile input.

11. The apparatus according to claim 1, wherein the first instance and the second instance of the tile include different numbers of inputs and different amounts of communication resources.

12. The apparatus according to claim 1, wherein the fabric circuit includes: A chain of tile instances including the first instance and the second instance of the tile, wherein adjacent tile instances in the chain are connected via at least a portion of the tile input and tile output of a given tile instance.

13. The apparatus according to claim 1, wherein the apparatus is a computing device, the computing device further including: A display; And A network interface circuit.

14. The apparatus according to claim 1, wherein the processor includes: A plurality of single instruction multiple data pipelines configured to execute instructions; And Fixed function circuitry configured to control the single instruction multiple data pipelines to perform operations for at least one of the following types of programs: Graphics shader programs; and Machine learning programs.

15. A method, the method comprising: A computing system assigns communication resources of a tiled fabric circuit for a given cycle, wherein the fabric includes at least a first instance and a second instance of a tile, wherein the tile includes: A client input configured to interface with client circuitry of the computing system; A tile input configured to interface with one or more other tile instances; and Communication resources that can be assigned to the client input and the tile input, wherein the communication resources include: A plurality of internal links; A client output configured to interface with client circuitry; and A tile output configured to interface with one or more other tile instances; wherein the assignment includes assigning communication resources of the tile instance to at least a portion of the client input and the tile input for a next cycle based on priority information of the input for a given tile instance; and updating, by the computing system over multiple cycles, the priority information for a given tile instance of the structured circuit based on the assignment result.

16. The method according to claim 15, wherein the assignment includes: determining priority information of the input for the tile instance; determining, at least partially in parallel for a plurality of communication resources and a plurality of inputs, the number of inputs having a higher priority than a given input and requesting the same resource; determining, for a given resource, the number of inputs that the resource can serve in the given cycle; and assigning communication resources to the inputs for the given cycle based on the determination of the number of inputs.

17. A non-transitory computer-readable storage medium storing design information, the design information specifying a design of at least a portion of a hardware integrated circuit in a format recognizable by a semiconductor manufacturing system, the semiconductor manufacturing system being configured to use the design information to produce the circuit according to the design, wherein the design information specifies that the circuit includes: a plurality of client circuits; a structured circuit, the structured circuit including at least a first instance and a second instance of a tile, wherein the tile includes: a client input configured to interface with a client circuit; a tile input configured to interface with one or more other tile instances; and communication resources that can be assigned to the client input and the tile input, wherein the communication resources include: a plurality of internal links; a client output configured to interface with a client circuit; and a tile output configured to interface with one or more other tile instances; a control circuit configured to: in a given cycle, assign communication resources of the tile instance to at least a portion of the client input and the tile input for a next cycle based on priority information of the input for a given tile instance; and update, over multiple cycles, the priority information for a given tile instance of the structured circuit based on the assignment result.

18. The non-transitory computer-readable storage medium according to claim 17, wherein, in order to assign communication resources of a given tile instance, the control circuit is configured to: determine priority information of the input for the tile instance; determine, at least partially in parallel for a plurality of communication resources and a plurality of inputs, the number of inputs having a higher priority than a given input and requesting the same resource; determine, for a given resource, the number of inputs that the resource can serve in the given cycle; and assign communication resources to the inputs for the given cycle based on the determination of the number of inputs.

19. The non-transitory computer-readable storage medium according to claim 17, wherein the communication resources include: dedicated resources for one or more non-blocking virtual channels; and Arbitration resources for a stoppable virtual channel.

20. The non-transitory computer-readable storage medium according to claim 17, wherein, in order to update the priority information for a given tile instance based on the assignment result over a plurality of cycles, the control circuit is configured to: classify the inputs to the tile instance based on whether all or a portion of the requested resources are received over the plurality of cycles for the inputs to the tile instance; and update the priority for the inputs to the tile instance based on the classification.

Citation Information

Patent Citations

  • System, Apparatus And Method For Multi-Die Distributed Memory Mapped Input / Output Support

    US20190303334A1

  • Partial write management in a multi-tiled compute engine

    US20210056028A1