Array processing entity array layout and routing

Through the fast local layouter QRAPP and Benes network wiring scheme, the processing entity layout and wiring in the entity array is automated, solving the time-consuming and inefficient problems in the prior art, and achieving an efficient and economical layout and wiring process.

CN120112899APending Publication Date: 2025-06-06MOBILEYE VISION TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380072513.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-13
Filing Date
2023-10-13
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is time-consuming and inefficient in processing entity layout and routing in solid arrays, lacking efficient, economical and scalable solutions.

Method used

An automated method is adopted to realize the rapid layout and routing of processing entities in the processing entity array through the fast local layout QRAPP and the Benes network wiring scheme. The method includes balancing the calculation graph, segmenting it into subgraphs, applying local and remote connectivity constraints, and optimizing the layout and routing process using hash values.

Benefits of technology

It significantly reduces the consumption of computing resources, storage resources and traffic, and realizes efficient automation of processing entity layout and routing in processing entity arrays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120112899A_ABST
    Figure CN120112899A_ABST
Patent Text Reader

Abstract

A method is provided for automatically placing and routing processing elements (PEs), where each PE has limited connectivity to its neighbors. The method may obtain layout primitive (PP) position window information that defines possible positions of a set of PPs within a coarse-grained reconfigurable array (CGRA). The method may obtain hardware constraints with respect to the PE array, the hardware constraints including local connectivity constraints and remote connectivity constraints. The method may receive a computational graph (CG) representing a mathematical expression to be computed by the PE array. The method may determine a location of the PE in the PE array based on the CG, the PP location window information, and hardware constraints.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] priority

[0002] This application claims the benefit of priority to U.S. Provisional Patent Application Serial No. 63 / 415,945, filed on October 13, 2022, which is incorporated herein by reference in its entirety. Background Art

[0003] The need to increase the throughput and performance of computerized systems has driven the industry to use arrays of processing entities. An array of processing entities may be designed to implement a computational graph. The placement and routing process for determining the locations of processing entities in the array and the connectivity between processing entities is time consuming and may take hours to complete. There is an increasing need to provide an efficient, economical, and scalable solution for the placement and routing of processing entities in an array of processing entities. Summary of the invention

[0004] A method and a non-transitory computer readable medium for placement and routing of processing entities in an array of processing entities may be provided. BRIEF DESCRIPTION OF THE DRAWINGS

[0005] In the drawings, which are not necessarily drawn to scale, like numbers may describe similar components in different views. Like numbers with different letter suffixes may represent similar different instances. In the various figures of the accompanying drawings, some embodiments are illustrated by way of example and not by way of limitation, in which:

[0006] Figure 1 An example of processing an array of entities is illustrated;

[0007] Figure 2 An example of a computational graph is illustrated;

[0008] Figure 3 An example of the balancing step is illustrated;

[0009] Figure 4 An example of generating a subgraph is illustrated;

[0010] Figure 5 An example of layout primitives is illustrated;

[0011] Figure 6 An example of layout primitives is illustrated;

[0012] Figure 7 Examples of allowed hardware implementations of layout primitives and permutations are illustrated;

[0013] Figure 8 An example of the method is illustrated;

[0014] Fig. 9 An example of the method is illustrated;

[0015] Fig.10 An example of the method is illustrated;

[0016] Fig.11 An example of the method is illustrated;

[0017] Fig.12 An example of layout primitives and generation of consumer hash values ​​is illustrated;

[0018] Fig.13 An example of layout primitives and generation of consumer hash values ​​is illustrated;

[0019] Fig.14 Examples of consumer hash values ​​for layout primitives and allowable hardware implementations of the layout primitives are illustrated;

[0020] Fig.15 Examples of consumer hash values ​​for layout primitives and allowable hardware implementations of the layout primitives are illustrated;

[0021] Fig.16 Examples of allowable hardware implementations of layout primitives and generation of consumer hash values ​​are illustrated;

[0022] Fig.17 Examples of producer hash values ​​for layout primitives and allowable hardware implementations of the layout primitives are illustrated;

[0023] Fig.18 An example of a producer hash value for an allowable hardware implementation of a layout primitive is illustrated;

[0024] Fig.19 Examples of windows, anchors, layout primitives, and two layout iterations are illustrated;

[0025] Fig. 20 Examples of windows and hash values ​​are illustrated;

[0026] Fig.21 Examples of layout primitives and full hash values ​​are illustrated;

[0027] Fig. 22 An example of the method is illustrated;

[0028] Fig.23 An example of the method is illustrated;

[0029] Fig.24 An array of PEs is instantiated; and

[0030] Fig.25 An array of PEs is illustrated. DETAILED DESCRIPTION

[0031] In the following detailed description, numerous specific details are set forth in order to provide a thorough understanding of the embodiments of the present disclosure. However, it will be appreciated by those skilled in the art that the present embodiments of the present disclosure may be practiced without these specific details. In other cases, well-known methods, procedures, and components are not described in detail to avoid obscuring the present embodiments of the present disclosure.

[0032] Provided herein is a method for automatically placing and routing in only a few seconds to a few minutes, as opposed to prior art methods that take longer (e.g., by factors exceeding x10, x20, x30, and even more). This provides a significant reduction in computational resources (e.g., by factors exceeding x10, x20, x30, and even more) and / or a significant reduction in memory resources (e.g., by factors exceeding x10, x20, x30, and even more) and / or a significant reduction in traffic (e.g., by factors exceeding x10, x20, x30, and even more).

[0033] An array, such as a coarse-grained reconfigurable array (CGRA) of processing elements (PEs), also called processing entities, can be abstractly viewed as a lattice of processing elements (PEs), where each PE has limited connectivity to its neighbors, a feature known as locality constraints.

[0034] For each PE, there are two kinds of locality-constrained neighbors: a set of local data producers, those neighboring PEs to which data can be transferred; and a set of local consumers, defined similarly. See, for example, Figure 1 - CGRA 10 includes multiple PEs, such as a certain PE 11 having eight neighbors - producer 12, consumer 13 and PE 14 which is both a consumer and a producer.

[0035] A CGRA implementation may allow a set of additional shallow connectivity to non-neighbor PEs, which is often small and insufficient to support complex computations.

[0036] A computational graph represents the input program, where different nodes represent operations and (directed) edges represent producer-consumer relationships.

[0037] In this computational graph, PEs are represented by nodes (solid circles). For example, see Figure 2 , which illustrates the calculation of the function F = (A + B) * (B + C) Fig. 20 , the dashed circle 21 illustrates that variables A, B and C are provided to two adders (nodes 22 and 23), which are two PEs that perform addition, thereby feeding their outputs to a multiplier (node ​​24).

[0038] Local connectivity constraints are modeled as edges between PE nodes.

[0039] CGRA may be required to implement the computation graph.

[0040] The computation graph may be balanced and then partitioned into subgraphs.

[0041] Partition points (links between subgraphs) can be routed to Benes Networks (BEs).

[0042] Each such subgraph is considered an atomic unit that should be laid out via local connectivity.

[0043] One or more subgraphs should be located within a 'bucket' or window.

[0044] The desired number and layout of buckets can be predefined, while allowing edges (ie, computation dependencies) to switch between local or network connections within a subgraph if necessary to maintain timing validity.

[0045] Each bucket is then fed into a fast local placer, called QRAPP (Quite Fast Affine Placement Planner), which quickly answers whether a local placement is possible (and returns such a placement if possible). As described herein, the entire computational flow may include balancing, relaxing, bucketing, placement, undoing, and routing. QRAPP may be used during the placement phase of this overall computational flow.

[0046] Once placement for all buckets is achieved, routing is performed to allow both broad expressivity and an efficient and fast routing process.

[0047] Assuming that PE-to-PE connectivity is local, the computation graph can be balanced. Balancing can equalize the delays between different paths between the same pair of source and destination of the computation graph.

[0048] The balancing may be followed by a relaxation phase. This may include identifying graph instances that have computationally complex topologies, topologies that are not feasible for the target hardware, or topologies that may cause stress to bucketing or layout activities. Once these graph instances are identified, relaxation includes reallocating these graph instances into subgraphs. In one example, if the identification phase identifies a graph instance node with six child nodes, the reallocation phase may include cloning the node and moving three of the six child nodes to the cloned node. The relaxation phase may provide improved processing throughput, such as by reducing overall compilation time.

[0049] For each bucket, this approach is repeated, for each subgraph that should be within the bucket: (a) try a local layout (by calling QRAPP), (b) if that fails, try to relax the local constraints by 'removing' local edges (turning the network) from the bucket's subgraph while maintaining timing correctness, or (c) if that fails, try another bucket layout.

[0050] QRAPP may identify a special predefined set of subgraphs in its input, decompose into their combinations, and utilize a pattern table for each such special subgraph in order to overcome the complexity of implementing which set of local connections are allowed and how to connect several such connections.

[0051] Once all buckets are laid out, an undo phase may be used to reduce or minimize the amount of remote edges caused by bucketing. This may include checking both remote edges within each bucket (e.g., local edges due to proximity) and remote edges across bucket edges (e.g., based on bucket proximity). Once checked, the undo may include reorganizing buckets across the PE grid to improve or maximize cross-edge opportunities. One or more optimization solutions such as constraint satisfaction problem (CSP) optimization, integer linear programming (ILP) optimization, or proposition satisfiability problem (SAT) optimization may be used to solve the determination of such cross-edge opportunities. Undo may also include ensuring that timing validity remains intact after such reorganization, and undoing any reorganization if timing validity does not remain intact.

[0052] This use of revocation provides various advantages. Because remote edges tend to carry increased latency (compared to local edges), reducing or minimizing the revocation process of remote edges can reduce or minimize computational latency. Additionally, removing remote edges during revocation can improve routing feasibility, such as by increasing the likelihood that the resulting net transition may be both feasible and quickly resolved.

[0053] Once all buckets are placed and undo is complete, the Benes Network (BN) routing scheme can be applied, which can routinely generate rich classes of transformations completely and simply (efficiently) through BN, and can also be applied to sub-transformations of an overall BN transformation (the method can declare routability by scanning non-intersecting sub-transformation parts of the overall transformation).

[0054] Routing may include using the freedom levels embedded in the PE array that represents the net (transition), shaping it into transitions that conform to the BN, and completing the routing.

[0055] In another example, the method may receive some input transformation and relax it in such a way as to make it 'as much' as ​​possible routable.

[0056] balance

[0057] The computation graph can be balanced to provide equal latency to different paths between the same source and the same destination.

[0058] In order for a node in a computation graph (which is a PE) to operate correctly, its inputs should arrive in the same cycle.

[0059] As an example, consider Figure 3 The following subgraph 40 in FIG. The subgraph 40 includes a source node 41 and a target node 46. There are two paths between the source node 41 and the target node 46.

[0060] If each edge does not contribute to latency (in cycles), then each PE contributes a minimum latency of 3 cycles, and may additionally delay the data by another 5 cycles.

[0061] Under this assumption, the left path 48 (including nodes 42, 43 and 44) ​​has a (minimum) delay of 9 cycles, while the right path 49 (including node 45) has a (maximum) delay of 8 cycles.

[0062] A potential balance would include adding 1 delay to the right path, such as adding a duplicate node 47 between nodes 41 and 45 or between nodes 45 and 46. By doing this, we increase the minimum delay of the right path to 6 cycles, and since each node may contribute up to 5 cycles of additional delay, for example, we can use 3 cycles of delay from duplicate node 47, and we have the right path delay and left path delay equalized at 9 cycles.

[0063] The same process should be applied to other paths of the computation graph.

[0064] This problem is considered as a non-polynomial (NP) problem and can be solved using ILP (integer linear programming), which is the NP derivation of LP (linear programming), where the problem to be solved is reconstructed as a set of linear equations, that is, the equation form is: Ax+By+Cz…=R, where A, B, C… are constant coefficients, and x, y, z,… are variables whose values ​​are sought.

[0065] The balance under the assumption is that any edge between two DPUs is a local edge, that is, an edge with zero additional delay.

[0066] Even if the assumption of zero latency is wrong, it can be remedied during the 'bucketing' phase, where the assumption can be broken to make the computation graph more feasible for layout reasons.

[0067] Bucketing

[0068] A balanced computational graph may be partitioned (e.g., bucketed) into disjoint subgraphs to speed up layout (laying out N subgraphs of maximum size M in series is much simpler than laying out a single graph of maximum size N x M), where N and M are positive integers.

[0069] The subgraphs are initially linked by edges, and the partitioning comprises replacing these edges with long-range connectivity links (eg with ports of a net or interconnect, such as a Benes net), for example by replacing the links with long-range connectivity links, without hardware routing constraints.

[0070] For example, see Figure 4 The computation graph 30 includes subgraphs 31, 32 and 33 linked by edge 35. Edge 35 is considered a long-range connectivity link and is "removed" to provide disjoint subgraphs.

[0071] The replacement should be done by taking into account remote connectivity constraints, such as limited remote link capacity at increased latency and / or bandwidth per PE (e.g., up to a certain number of bits per cycle - e.g., up to 32, 64, 128, etc. bits per cycle). In one example, the replacement may be determined based on a reduction or minimization of latency or remote link capacity to improve processing efficiency and bandwidth. Increased use of remote connectivity constraints may reduce the number of layout constraints, which may increase the layout stage and increase latency of the execution graph. In one example, the number of remote connectivity constraints may be selected to improve or maximize layout stage speed while improving or minimizing latency of the execution graph, such as using one or more optimization solutions.

[0072] No additional replication nodes (such as Figure 3 Node 47 in - these additional replica nodes simply copy their input to their output, while adding latency) may also be beneficial.

[0073] Therefore, the starting point of the problem is that: (a) the balanced computation graph (before partitioning) is balanced (under the assumption that all local point-to-point PE connections impose layout constraints (e.g., local edges have zero latency)); (b) in terms of placement, the partitioned computation graph should be partitioned into N disjoint subgraphs, each of which has a maximum size M; and (c) the remote connectivity constraints should not be violated.

[0074] This problem can be solved by constructing a set of linear integer equations (i.e., ILP form) that stipulates the following:

[0075] a. Each node (HC) can be associated with any subgraph from the N available subgraphs.

[0076] b. If two nodes connected by an edge are associated with different subgraphs, their connecting edge should be a distant edge.

[0077] c. If an edge can be marked as far, then even if its two endpoint nodes are on the same subgraph, this is allowed if it is necessary to achieve a feasible balance.

[0078] d. Balance and correctness should be maintained.

[0079] e. Each single node may not exceed the PE limit of incoming far connections per node.

[0080] f. The subgraph of each bucket may not exceed the maximum size M.

[0081] These statements model a valid state, and this valid state enables the linear integer equations to be solved.

[0082] Once a valid ILP solution is found, it provides a valid segmentation.

[0083] The valid partitions may be referred to as buckets, and various groups of layout primitives may be fitted into one or more of these buckets.

[0084] Multiple layout primitives may be determined in advance, for example to include all or some of the possible layout primitives given various computation graph shapes and / or computation subgraph shapes.

[0085] The layout primitives to evaluate may be determined based on the computation graph.

[0086] Layout primitives express local connectivity constraints.

[0087] Since local connectivity constraints tend to be limited, the variety of layout primitives evaluated with respect to the computation graph is expected to be proportionally limited as well.

[0088] An example of a layout primitive is the fully connected (bipartitioned) layout primitive. Every node (producer) on one partition is connected (via local connections) to all nodes (consumers) on the other partition. For example, see Figure 5 A layout primitive 50 comprising two producers (P1 51 and P2 52) and three consumers (C1 61, C2 62 and C3 63) that are fully connected.

[0089] It should be noted that the above layout primitives can also be viewed as a combination of two 1-3 layout primitives, or a combination of six 1-1 layout primitives, and so on.

[0090] A fully connected layout primitive may contain many (perhaps all in context) locality constraints that can be easily gleaned from the computation graph, and it does not represent a complex structure (mostly or only locally connected), so it is expected that hardware matching structural patterns (a subset of their possible locally connected embeddings) is reasonable.

[0091] Since the presence of edges is required for significant biclinicity, and since computational graphs may also contain isolated nodes, a fully connected (biclinic) layout primitive is chosen.

[0092] It should be noted that in this preliminary step, the layout primitives also allow for the determination of some infeasibility cases (e.g., can help determine local connectivity that is required for the computation graph but is not feasible on local connections alone).

[0093] The method may check only the largest layout primitive, where the largest layout primitive may be defined as a layout primitive whose size may not exceed M (the maximum size of a subgraph) or may not exceed a position window in which the layout primitive should be laid out. The position window may define possible positions for a set of layout primitives, and different position windows may be used for different sets of layout primitives. The position window size may be equal to the bucket size, or may be smaller than the bucket size.

[0094] About the maximum layout primitives, and refer to Figure 5 , 2 producers to 3 consumers layout primitive 50 includes smaller layout primitives, such as a single producer to three consumers, a single producer to two consumers, and a producer to one consumer. The upper limit on the size of the maximum layout primitive is set by the size of the window. For example, a 5×5 window can support larger layout primitives, such as a 2 producer and four consumer layout primitive (see Figure 6 Layout primitive 70), 4 producer and 5 consumer layout primitives, etc.

[0095] Using the largest layout primitive reduces the computational resources required to perform population because the largest layout primitives are fewer than smaller layout primitives, thus providing fewer overall options to traverse. In addition, using the largest layout primitives tends to have fewer structural alternatives (ie, patterns allowed by the HW).

[0096] The maximum layout primitives are defined per computation graph.

[0097] The evaluated layout primitives should cover the entire computation graph.

[0098] Among other things, for each window, the process computes the effective positions of one or more layout primitives within the window.

[0099] PEs may differ from each other in functionality and their possible valid locations, and the padding also responds to these constraints. For example, there may be general purpose PEs, some PEs with trigonometry calculations, some other PEs with integer division capabilities, some fixed point PEs, some floating point PEs, etc. The possible valid locations may be driven by the desired accessibility of PEs with certain functionality to other PEs that do not have this functionality.

[0100] The filling process proceeds to perform calculations based on layout primitives.

[0101] Each placement primitive is associated with a matching table (or another type of information) of allowed hardware implementations, given hardware locality constraints. Figure 6 The layout primitive 70 includes two producers (P1 and P2) and four consumers (C1, C2, C3, C4).

[0102] Figure 7 Six allowed hardware implementations 81-86 of the layout primitive 70 are illustrated, along with six permutations 91(1)-91(6) of the first allowed hardware implementation 81.

[0103] The allowed hardware implementations may be represented by one or more tables. Such tables are pre-generated and may not be computationally graph-related. These tables may be generated once and / or offline and / or during an initialization step, for example as part of the creation and / or construction of a layout tool, and may be immediately derived from local connectivity constraints. These tables may store possible allowed hardware implementations for each layout primitive (assuming an empty window), so it can and should be calculated independently of the computational subgraph and / or in an offline state. By avoiding the use of computational subgraphs, the hardware implementation for each layout primitive can be determined more efficiently. Avoiding the use of CGs when determining PP locations provides various advantages. Using these predefined tables can save a lot of wasted computation cycles, especially compared to solutions that include computing allowed hardware implementations during the layout process. These tables can be stored in a compressed (e.g., compact) form. In one example, this may include representing one or more of these tables with one or more hash values. This compressed or compact storage format may include offline generation of tables, which allows for more complex methods that can be used to generate compressed forms of these tables.

[0104] With respect to layout primitives, the allowed hardware implementations of the layout primitives may be hashed (to provide a hash table accessed using the hash value) and ordered (in the hash table) to allow fast and simple access. For example, when searching to place a layout primitive in a window already populated with another layout primitive, rather than evaluating all possible allowed hardware implementations, the search may be limited to allowed hardware implementations having hash values ​​equal to (or at least similar to) an already placed anchor.

[0105] When hashing, the same index function may be used to hash the current layout primitive state to implement at an early stage whether the current layout primitive state can be laid out into a partially filled window.

[0106] The hashing can be producer-based, and / or consumer-based, and / or free-layout-based, etc.

[0107] The hash may represent a canonical format of a layout primitive (which may be an aligned format of the layout primitive), and may represent a non-canonical format of the layout primitive. The latter also provides an indication of the location of the layout primitive (or a node of the layout primitive) within the window.

[0108] Hash values ​​can be used to determine whether the current layout primitive can be positioned within an already partially filled window, for example by applying logical functions (e.g., NAND, AND, etc.) to various hash values, while saving computational resources because hash values ​​can provide a compact representation that allows multi-dimensional hardware implementation.

[0109] Using hashing and applying a logical function allows the amount and location and semantics of anchors (nodes of a layout that have been positioned layout primitives) to be easily determined.

[0110] Using a hash and applying a logic function provides a balanced tradeoff between fast and accurate access to the allowed hardware implementation and the amount of information required to store the hash value.

[0111] The hash value of the allowed hardware implementation can be generated by scanning the nodes of the allowed hardware implementation (e.g., starting from a predefined element of the window and ending at the end of the allowed hardware implementation). Alternatively, the scan can start from the end of the allowed hardware implementation and can end at a predefined point.

[0112] Here is an example of hashing pseudocode:

[0113] for(auto point:points){

[0114] unsigned adj_row=point.first-min_row;

[0115] unsigned adj_col=point.second-min_col;

[0116] bitmap_key|=1<<(adj_row*COL_DIM+adj_col);

[0117] unsigned adj_row2=point.first-p_min_row;

[0118] unsigned adj_col2=point.second-p_min_col;

[0119] key|=1<<(adj_row2*COL_DIM+adj_col2);

[0120] }

[0121] in:

[0122] points: A list of nodes in the primitive represented by (row, column) coordinates that can be producers or consumers or any other subgroup of nodes.

[0123] min_row / min_col: minimum row / column of the coordinates of the nodes in the 'points' list

[0124] p_min_col / p_min_row: minimum row / column among the coordinates of all nodes in the primitive

[0125] COL_DIM: is the size of the column size.

[0126] Next, construct the layout primitive layout plan.

[0127] The general orientation is to minimize the degree of freedom for each subsequent layout primitive being placed, primarily by following node / edge overlap.

[0128] The plan is then traversed in order, and any redundant layout primitives (layout primitives whose all nodes [singleton] and edges are already embedded by previous layout primitives) are removed.

[0129] Each such computation may have various overheads, and they may aggregate into a huge cost. By using static planning, no dynamic selection is required, which has been found to result in excellent running times.

[0130] Figure 8 An example of a layout process 800 is illustrated.

[0131] The layout process may begin by setting (802) a control variable K to a particular value (eg, 0).

[0132] The layout process is followed by an ordered list of hypothetical layout primitives, evaluating (804) the Kth layout primitive, ie, the layout primitive at the Kth position.

[0133] Step 804 may include checking whether the Kth layout primitive has an anchor point (i.e., a node that has been laid out in the window), if so, laying out the Kth layout primitive so that the nodes of the Kth layout primitive are laid out on the anchor point, if not, the process may use the layout order of the layout primitives. Both anchors in the case of an anchor point and the node to be laid out first (bootstrap) in the case of an anchor point are predetermined during the planning phase. In addition, a hash may be used to represent the anchor point.

[0134] Step 804 may be followed by step 806 which determines how to proceed (eg, update the value of K and jump to step 804), ending the process if the process succeeded or failed.

[0135] Steps 804 and 806 may be repeated multiple times during multiple iterations, and a position is selected from the possible valid positions (after intersecting it with the current idle frame position). This may also include finding an allowed hardware implementation that matches the current anchor relative position, which is also feasible after intersecting with the current idle position on our frame.

[0136] Assuming no anchor point exists, the relevant window of the layout frame is constructed and intersected with the candidate pattern.

[0137] If no pattern passes the intersection, a retreat action may be performed to retreat to K-1, otherwise the pattern is committed and nodes are arranged on it, and the value of K is increased, causing a push action.

[0138] In case of a push action, go to step 804.

[0139] In the case of a retreat action, then the simple flow for retreat is to permutate until no more permutations are available, then advance to the next mode.

[0140] The retreat may consist of achieving an intersection between the current layout primitive nodes (which are added to the framework by this retreat (denoted A)) and the set of nodes from the entirety of the layout primitive we have retreated from. If they intersect, A is further permuted (if possible), and re-advanced to K+1.

[0141] If no more permutations are available, advance to the next pattern. If no more patterns are available, retreat to K-1.

[0142] If A is empty, advance to the next pattern, permute, and go to K+1.

[0143] If the retreat continues until K=-1, it is declared infeasible.

[0144] If we advance until K = size(plan_list), (plan_list is the number of PPs in the layout plan), we have a valid solution and thus terminate successfully.

[0145] It should be noted that the inference of A can and should be done statically (once) at planning time.

[0146] Regarding ranking, it is important to distinguish logical groups within the pattern, taking a bipartite cluster as an example, which is two disjoint groups, producers and consumers, each with its own semantics. Producers can only be ranked with another producer, and consumers can only be ranked with another consumer. In general, each node can only be ranked with other nodes of its category.

[0147] Fig. 9 An example of method 400 is illustrated.

[0148] Method 400 may include an initialization step 410 of obtaining PE array hardware constraints.

[0149] Step 410 may include obtaining layout primitives and allowed hardware implementations of the layout primitives.

[0150] The method 400 may also include a step 420 of receiving a computation graph to be implemented by the PE array.

[0151] Step 420 may be followed by step 430 of performing one or more computation graph related operations. The computation graph related operations include generating a subgraph.

[0152] The generation of the subgraphs may be preceded by a balanced computational graph. The generation of the subgraphs may include partitioning the graph into subgraphs while maintaining balance, replacing links that previously linked the subgraphs with long-range connectivity while taking into account long-range connectivity constraints. This may involve using ILP.

[0153] Step 430 may include ordering layout primitives associated with the computation graph.

[0154] The order of primitives may be determined in any manner, for example, it may be determined based on at least one of the following constraints:

[0155] a. The number of nodes in the layout.

[0156] b. If there is no node for the placement, calculate the bootstrap score.

[0157] c. Reduce the number of permutations.

[0158] d. The maximum number of nodes reachable from a consumer.

[0159] e. Prefer fully laid out producer / consumer.

[0160] f. Minimum primitive size.

[0161] g. Producer of maximum layout.

[0162] h. Consumers of the largest layout.

[0163] When no anchor point is present, a guidance score is assigned. A node's guidance score is high when (a) there are few nodes reachable from the node, and / or (b) there are many nodes reachable. For each primitive, the guidance score is the same as the guidance score of the higher node. Moreover, such a node is a guidance for that primitive.

[0164] Once the order is set, the layout process can ignore layout primitives whose all nodes appear in a previous layout primitive and whose all edges appear in a previous layout primitive.

[0165] The layout process can also ignore infeasible locations within the window. For example, assuming that nodes at the corners of the window are restricted to connecting to two or fewer nodes, a source node connected to three or more destination nodes cannot be placed in the corners of the window.

[0166] Step 430 may be followed by step 440 of performing sub-image processing. Step 440 may include attempting to fit a position primitive representing a portion of the sub-image to the window.

[0167] Depending on the window, step 440 may include:

[0168] a. Calculate the effective position for each node.

[0169] b. Try to lay out the layout primitives.

[0170] i. If successful, proceed to the next layout primitive.

[0171] ii. If that fails, try the next layout option for that layout primitive.

[0172] iii. If all options fail, fall back to the previous layout primitive.

[0173] iv. Succeeds if all layout primitives representing the subgraph are laid out, otherwise fails.

[0174] Given a layout primitive to lay out.

[0175] a. If the layout primitive does not have a node to layout, then layout the guide node.

[0176] b. Try all free valid locations for the boot node.

[0177] c. Based on the nodes and free places of the layout within the window, try all valid solutions for the layout primitive.

[0178] d. If there is more than one unplaced producer / consumer.

[0179] e. For each of the chosen solutions, try all permutation options.

[0180] The planning phase is static, ie once the order is set it is fixed and will not change during the layout phase.

[0181] Planning can be done based on a given graph only, regardless of the hardware available elements, i.e., locations. In such cases, planning is performed once for all given groups of locations (i.e., various layout frames).

[0182] Another approach would be to perform planning for each set of positions, if planning can help pick a better ordering of layout primitives based on additional information, such as the valid positions of each node.

[0183] The goal of planning is to reduce option exploration so that if trying a subtree fails, the graph is tried as quickly as possible.

[0184] For example, layout primitives with anchor points are prioritized, and layout primitives with fewer permutations are preferred over those with higher permutations, etc.

[0185] The placeplanner may not be able to place a given graph into a given group of locations, so in this case it may be desirable to try placing the graph into a different group of locations.

[0186] Each group can be presented as a matrix where some entries are invalid.

[0187] A user may invalidate an entry because it is occupied by another computational graph, or because its corresponding hardware node does not implement the operation represented by the graph node for which the valid position matrix was generated.

[0188] For each node, a valid position matrix is ​​created, which can be used to determine whether a position is valid or invalid for such node.

[0189] In addition to the invalid positions passed by the user, more invalid positions may be determined based on the graph and hardware constraints, for example, node degree may affect whether a node can be placed in a matrix corner.

[0190] Such valid position matrices can be used to make retreat actions earlier and prevent the exploration of such failed options, identifying only after several layout primitives that there is no solution in such directions.

[0191] Hashes are used to index tables where the hardware implementation allows for storage.

[0192] For PE arrays, where the layout primitives are viewed as a bipartite group of producer nodes and consumer nodes, hashing can be performed by considering only producers and ignoring consumers, or vice versa, that is, considering consumers and ignoring producers, or considering both groups of nodes for hashing. Alternatively, the hashing process can consider producers, consumers, and idle nodes.

[0193] For example, a bitmap (or hexagonal string) may be used to hash matching patterns based on the positioning of considered nodes within the pattern, with each position assigned one or more bits.

[0194] In addition, the hash can be standardized to eliminate repeated patterns and offset from the matching position. To this end, the hash may not include the start of non-free nodes, for example, empty window edges (empty top or bottom rows and / or right or left columns) may be ignored.

[0195] For example, assuming the matching position is a rectangle where the left column or top row is empty, the pattern can be shifted left and up before computing the hash to produce a canonical hash. A canonical hash is a hash of the canonical format of the allowed hardware implementation. The hash can be applied to non-canonical formats of the allowed hardware implementation.

[0196] Once the canonical hash is created, the list of matching patterns can be reordered based on that hash. Because we can hash using only producers or only consumers or a mix of both or a mix of idle nodes, producers and consumers.

[0197] Since there may be more than one ordering that may be statically arranged and dynamically used based on the available set of nodes for hashing, the process may save only an ordered list of indexes to the original table of matching patterns, rather than saving multiple differently ordered lists of complete matching patterns.

[0198] Hashes are not only used to match patterns, but can also be used on subgroups of hardware locations represented as e.g. a matrix to capture free locations.

[0199] Hashes can be used to speed up the intersection of matching patterns with free positions. A unique technique is to precompute a dynamic hash of the free positions, which allows the hashes for several matrices near the anchor node to be regenerated using only simple arithmetic operations such as shifts and masks.

[0200] This may reduce the computation of the hash per matching pattern and thus enhance matching pattern intersection with free locations.

[0201] Fig.10 An example of a method 500 for processing placement and routing of physical (PE) devices in an array of PEs is illustrated.

[0202] Method 500 may include steps 510 and 520 .

[0203] Step 510 may include obtaining layout primitive (PP) position information regarding possible positions of a set of PPs within a window.

[0204] Step 520 may include obtaining hardware constraints on the PE array, which may include local connectivity constraints and remote connectivity constraints.

[0205] Local connectivity constraints may define point-to-point connections between PEs, where the point-to-point connections connect the PE to the neighbors of the PE. Local may refer to connections whose length is below a certain length threshold (e.g., the distance between adjacent PEs). In one example, the determination of the location of the PE may be based on the reduction or minimization of the length of the local connections in order to improve processing efficiency and bandwidth. Layout primitives may provide various advantages, such as improving or maximizing comprehensive layout searches. Layout primitives may also be used to improve or optimize the compactness of a layout, such as by allowing only compressed forms of layout primitives.

[0206] For example, refer to Fig.25 , the PE may include Q1 local input ports 511(1)-511(Q1) and Q1 output ports 512(1)-512(Q1) for local communication with Q1 neighbors, and may include I / O ports for remote communication (RC) (e.g., with one or more Benes networks). The I / O ports for RC may include Q2 input ports 513(1)-513(Q2) and Q3 output ports 514(1)-514(Q3) (e.g., 1 unicast BN output port and 2 unicast BN input ports, but other numbers may be provided). The PE may include Q4 input ports 515(1)-515(Q4) for receiving broadcasts (provided simultaneously to a group of PEs, such as a row of PEs or a column of PEs in the array or any 2D group of PEs). Q1, Q2, Q3, Q4, and Q5 are positive integers that may exceed two. For example, Q1 may be two or greater, Q1 may range between 4 and 12, Q1 may be equal to 8, etc.

[0207] Remote connectivity constraints may limit the number of remote connections of a single PE.

[0208] The hardware constraints may include different functions of at least two PEs in the PE array.

[0209] Steps 510 and 520 may be followed by a step 530 of receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array.

[0210] The PP location information may be determined independently of the CG. For example, step 510 may be calculated prior to receiving the CG, may be calculated offline, etc. By avoiding the use of computational subgraphs, the hardware implementation for each layout primitive may be determined more efficiently. [This may be the same added text we used above for paragraph

[0092] . We hope that both locations describe the improvements provided by avoiding the use of the CG to determine the PP location.]

[0211] Step 530 may be followed by step 540 of determining the location of the PEs in the PE array based on the computation graph, the PP location information, and the hardware constraints.

[0212] In some computation graphs, there may be consumers that are not linked to producers. This can be solved by using PP to represent consumers.

[0213] Generally, a PP may represent connectivity between one or more producer PEs and one or more consumer PEs.

[0214] The PP may be a fully connected PP.

[0215] Step 540 may include at least one of the following steps 540(a) to 540(t):

[0216] a. Split the CG into sub-images.

[0217] b. Balance CG.

[0218] c. Use integer linear programming (ILP) to partition the CG into subgraphs.

[0219] d. Segment the CG using any computational process that does not include ILP.

[0220] e. Split the CG into subgraphs and define the edges between the subgraphs as long-range connectivity connections. Long-range connectivity connections can be connections to the Benes network.

[0221] f. Compensate for the delay introduced by the remote connectivity connection.

[0222] g. Fit subgraphs to buckets.

[0223] h. Fit the PP to a window. The window may be equal to the bucket, or may be different from the bucket. The size of the window may exceed the corresponding size of the bucket. The buckets may include at least two buckets of different shapes from each other. The buckets may include at least two buckets of different sizes from each other.

[0224] i. Apply multiple layout iterations.

[0225] j. Apply multiple layout iterations based on the PP position information.

[0226] k. Perform multiple sets of layout iterations.

[0227] l. Check the position of PP.

[0228] m. Check the arrangement of PPs given one or more positions of PPs.

[0229] n. Performing multiple sets of layout iterations, wherein a set of layout iterations may include: (a) evaluating gradually increasing combinations of PPs during different layout iterations of the set until a combination is found to be potentially infeasible; (b) gradually reducing the combination;

[0230] and (c) gradually increasing the combination.

[0231] o. Perform multiple sets of layout iterations, where a set of layout iterations may include evaluating increasingly larger combinations of PPs during different layout iterations of the set until all subgraphs are laid out in buckets.

[0232] p. Perform multiple groups of layout iterations, wherein a group of layout iterations may include: (a) selecting a PP to be evaluated according to an order of the PPs to be evaluated to provide a selected PP, (b) determining the feasibility of positioning the selected PP within the bucket and the position of the selected PP within the bucket is feasible, wherein the determination may be based on the PP position information and based on the current state of the bucket, the current state of the bucket may include s nodes of the bucket that may have been placed in the bucket, and (c) when there is a feasibility of positioning the selected PP within the bucket, jump to step (a).

[0233] q. Perform multiple sets of placement iterations, where a set of placement iterations may include checking the feasibility of placement of a combination of PPs based on non-feasible placement information.

[0234] r. Check the combination of PP positions.

[0235] s. Check the permutations and combinations of PP.

[0236] t. Determine the order of PPs to be evaluated during multiple layout iterations.

[0237] Step 540 may be followed by step 550 of forming a PE array by placing the PEs based on the determined locations of the PEs.

[0238] Fig.11 An example of a method 600 for processing placement and routing of physical (PE) devices in an array of PEs is illustrated.

[0239] Step 600 may include step 520 of obtaining hardware constraints on the PE array, which may include local connectivity constraints and remote connectivity constraints.

[0240] Step 520 may be followed by step 530 of receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array.

[0241] Step 530 may be followed by step 640 of making a placement primitive (PP) based determination of the location of the PEs in the PE array based on the computation graph, the placement primitives (PPs), and the hardware constraints. The PPs to be evaluated during step 640 are sorted. Step 640 may include sorting the PPs. The sorting may be determined once and then maintained during multiple fitting iterations.

[0242] Step 640 may be performed without receiving PP position information regarding possible positions of a set of PPs within the window. Alternatively, method 600 may include receiving PP position information and performing step 640 based at least in part on the PP position information.

[0243] Step 640 may include at least one of the following steps 640(a) to 640(s):

[0244] a. Split the CG into sub-images.

[0245] b. Balance CG.

[0246] c. Use integer linear programming (ILP) to partition the CG into subgraphs.

[0247] d. Segment the CG using any computational process that does not include ILP.

[0248] e. Split the CG into subgraphs and define the edges between the subgraphs as long-range connectivity connections. Long-range connectivity connections can be connections to the Benes network.

[0249] f. Compensate for the delay introduced by the remote connectivity connection.

[0250] g. Fit subgraphs to buckets.

[0251] h. Fit the PP to a window. The window may be equal to the bucket, or may be different from the bucket. The size of the window may exceed the corresponding size of the bucket. The buckets may include at least two buckets of different shapes from each other. The buckets may include at least two buckets of different sizes from each other.

[0252] i. Apply multiple layout iterations.

[0253] j. Apply multiple layout iterations based on the PP position information.

[0254] k. Perform multiple sets of layout iterations.

[0255] l. Check the position of PP.

[0256] m. Check the arrangement of PPs given one or more positions of PPs.

[0257] n. Performing multiple sets of layout iterations, wherein a set of layout iterations may include: (a) evaluating increasing combinations of PPs during different layout iterations of the set until a combination is found to be potentially infeasible; (b) gradually reducing the combination; and (c) gradually increasing the combination.

[0258] o. Perform multiple sets of layout iterations, where a set of layout iterations may include evaluating increasingly larger combinations of PPs during different layout iterations of the set until all subgraphs are laid out in buckets.

[0259] p. Perform multiple groups of layout iterations, wherein a group of layout iterations may include: (a) selecting a PP to be evaluated according to an order of the PPs to be evaluated to provide a selected PP, (b) determining the feasibility of positioning the selected PP within the bucket and the position of the selected PP within the bucket is feasible, wherein the determination may be based on the PP position information and based on the current state of the bucket, the current state of the bucket may include s nodes of the bucket that may have been placed in the bucket, and (c) when there is a feasibility of positioning the selected PP within the bucket, jump to step (a).

[0260] q. Perform multiple sets of placement iterations, where a set of placement iterations may include checking the feasibility of placement of a combination of PPs based on non-feasible placement information.

[0261] r. Check the combination of PP positions.

[0262] s. Check the permutations and combinations of PP.

[0263] Step 640 may be followed by step 650 of forming a PE array by placing the PEs based on the determined locations of the PEs.

[0264] In the following examples, hash values ​​may be provided in binary format and / or in hexadecimal format. Other formats may be used.

[0265] Fig.12 An example of a PP 100 is illustrated, which includes two producers (node ​​A and node B), collectively referred to as 100-1, and three consumers (node ​​C, node D, and node E), collectively referred to as 100-2. Fig.12 Also illustrated are an allowed hardware implementation 102 of the PP 100, a consumer table 103 showing only the consumer 100-2, and a canonical consumer table 103 showing only the consumer in canonical form, in this case aligned to the upper left corner of the table.

[0266] Fig.12 Also illustrated is a scan table 105 showing the generation of consumer hash values ​​obtained when scanning the consumer table 103 from the last consumer to the first consumer (in a raster scan pattern) and assigning set bits for consumers and zero bits for nodes not populated by a consumer.

[0267] Fig.13 PP (in Fig.12 100), a consumer table 113 showing only consumer 100-2, and a canonical consumer table 114 showing only the consumers in canonical form, in this case aligned with the upper left corner of the table.

[0268] Fig.13Also illustrated is a scan table 115 showing the generation of consumer hash values ​​obtained when scanning the consumer table 113 from the last consumer to the first consumer (in a raster scan pattern) and assigning set bits for consumers and zero bits for nodes not populated by a consumer.

[0269] Fig.14 An example of a PP 100 is illustrated, which includes two producers (node ​​A and node B), collectively referred to as 100-1, and three consumers (node ​​C, node D, and node E), collectively referred to as 100-2. Fig.14 Consumer hash values ​​for various allowed hardware implementations 121, 122, 123, 124, 125, and 126 of the PP 100 are also illustrated.

[0270] Fig.15 An example of a PP 101 is illustrated, which includes three producers (node ​​C, node D, and node E), collectively referred to as 101 - 1 , and a consumer (node ​​F), collectively referred to as 101 - 2 . Fig.15 Consumer hash values ​​for various allowed hardware implementations 171, 172, 173, 174, 175, and 176 of PP 101 are also illustrated.

[0271] Fig.16 An example of a PP 101 is illustrated, which includes three producers (node ​​C, node D, and node E), collectively referred to as 101 - 1 , and a consumer (node ​​F), collectively referred to as 101 - 2 . Fig.16 Also illustrated are allowed hardware implementations 171 and 179 of PP 101, producer tables 171' and 179' showing only producer 101-2, and a scan table 178 showing producer hash values ​​obtained when scanning the producer tables 171' and 179' from the last producer to the first producer (in a raster scan pattern) and assigning set bits to producers and zero bits to nodes not filled by producers.

[0272] Fig.17 An example of a PP 101 is illustrated, which includes three producers (node ​​C, node D, and node E), collectively referred to as 101 - 1 , and a consumer (node ​​F), collectively referred to as 101 - 2 . Fig.17 Also illustrated are allowed hardware implementations 131, 132, 133, 134, 135, and 136, along with their producer hash values, which are canonical hash values.

[0273] Fig.18Allowed hardware implementations 131, 132, 133, 134, 135, and 136 are illustrated, along with their producer hash values, which are non-canonical hash values ​​and include information about the location of the producer within the table. For example, the binary hash value of each of the allowed hardware implementations 134 and 136 ends with a reset bit to indicate that the top left corner node of the window is free.

[0274] Fig.19 A window 140 comprising four by six elements is illustrated. The window is implemented by the hardware enabled by the PP 100 ( Fig.14 126), which includes a first producer (node ​​A) located above a second producer (node ​​B), which is located above three consumers (nodes C, D and E), which are located side by side in the same row of window 140.

[0275] Window 140 is partially filled, and the layout process may include evaluating whether the next layout primitive (PP 101 ) may also fill window 140 .

[0276] Nodes C, D and E are already placed in the window (they are consumers of PP 100). They also belong (as producers) to the next placement primitive.

[0277] To reduce computation time, the placement process may select, from among multiple allowed hardware implementations of PP 101, only allowed hardware implementations having hash values ​​equal to the hash values ​​of already placed PP 100 (e.g., producer hash values, since nodes C, D, and E are producers of PP 101), namely, allowed hardware implementation 131 and allowed hardware implementation 132.

[0278] Fig. 20 An example is illustrated of evaluating multiple candidate windows 141, 142, and 143 within larger window 104. The candidate windows are shifted vertically relative to each other, and the layout process may determine in which candidate window, if any, an allowed hardware implementation of PP 101 may be located.

[0279] The per-window determination may include computing any type of hash value (consumer hash value, producer hash value, idle hash value, or full hash value) and performing one or more Boolean operations to determine whether an allowed hardware implementation of PP 101 may be located in any of the candidate windows.

[0280] exist Fig. 20 The allowed hardware-implemented producer binary hash values ​​(in Fig.14, represented as 126 in , is 111, and a NAND operation is applied between the producer binary hash value and the producer hash value of each of the first, second, and third candidate windows to indicate that only the third window is a valid window.

[0281] The next step may include making such a comparison to see if the consumer node (node ​​F) can be placed in the third window.

[0282] Fig.21 An example of a hash based complete map is illustrated, where a single hash value of an allowed hardware implementation represents consumers, producers, and vacancies of the allowed hardware implementation. The map associates a first value (e.g., 10) with each producer, a second value (e.g., 01) with each consumer, and a third value (e.g., 00) with each idle node. The map can be a canonical map or a non-canonical map.

[0283] The hash value may be obtained by scanning the allowed hardware implementations, eg, starting from the last non-free element of the allowed hardware implementations and going backwards, eg, until the first non-free element of the allowed hardware implementations or until a predefined position in a table.

[0284] The scanning may be according to a raster scanning pattern and may be performed in any direction (backward or forward).

[0285] Fig.21 An example of a PP 101 is illustrated, which includes three producers (node ​​C, node D, and node E), collectively referred to as 101 - 1 , and a consumer (node ​​F), collectively referred to as 101 - 2 . Fig.17 Allowed hardware implementations 151, 152, 153, 154, 155, and 156 are also illustrated, along with their full hash values.

[0286] Fig. 22 An example of a method 2200 for processing a hash-based layout of an array of physical entities (PEs) is illustrated.

[0287] The method 2200 may include a step 620 of obtaining hardware constraints on the array, the hardware constraints including local connectivity constraints and remote connectivity constraints.

[0288] The method 2200 may include a step 630 of receiving a computation graph (CG) representing a mathematical expression to be computed by the array.

[0289] Steps 620 and 630 may be followed by step 2230 of making a hash-based PP-based determination of the location of the PEs in the array based on the computation graph, placement primitives (PP), and hardware constraints.

[0290] Step 2230 may be followed by step 650 of forming a PE array by placing the PEs based on the determined locations of the PEs.

[0291] Step 2230 may include performing multiple sets of hash-based layout iterations, wherein a set of hash-based layout iterations is associated with a portion of the array.

[0292] Fig.23 A method 2300 for performing a set of hash-based layout iterations is illustrated.

[0293] The method 2300 may include a step 2310 of performing an initial hash-based layout iteration for initial filling of empty windows with initial allowed hardware implementations of an initial PP.

[0294] Step 2310 may be followed by step 2330 of performing additional hash-based placement iterations for filling the partially filled window with one or more additional enabled hardware implementations of one or more additional PPs.

[0295] Step 2330 may be performed based on one or more hash values ​​of one or more populated nodes.

[0296] Step 2330 may also be performed based on one or more hash values ​​implemented by additional enabled hardware.

[0297] Step 2330 may include multiple repetitions of step 2332 performing additional hash-based layout iterations. The repetitions may continue until a stopping condition is reached, such as exhausting the search, filling the window so that other allowed hardware implementations cannot be added to the window, etc.

[0298] Step 2332 may include a step 2334 of selecting an additional allowed hardware implementation from a set of additional allowed hardware implementations.

[0299] The selection may be based on (i) one or more hash values ​​of the members of the group and (ii) one or more hash values ​​of one or more populated nodes that appear at least in part among the members of the group.

[0300] The one or more hash values ​​of the members of the group may include a producer hash value and a consumer hash value.

[0301] The one or more hash values ​​of the members of the group may include a producer hash value, a consumer hash value, and an idle hash value.

[0302] The one or more hash values ​​for members of the group may include hash values ​​that may indicate producers, consumers, and idle nodes.

[0303] One or more hash values ​​of members of the group may be canonical hash values.

[0304] One or more hash values ​​of members of the group may be non-canonical hash values.

[0305] The selecting may include selecting a member of the group that has one or more hash values ​​that are equal to one or more hash values ​​of one or more populated nodes that appear in the members of the group.

[0306] The one or more hash values ​​of the members of the group may be one or more bitmaps.

[0307] Step 2332 may also include step 2336 (which follows step 2334) of determining whether additional enabled hardware implementations may be placed in the window that may be partially filled.

[0308] Step 2336 may include performing one or more Boolean operations.

[0309] The Boolean operation of the one or more Boolean operations may be applied to the hash value of the one or more populated nodes and to the hash value of the additional enabled hardware implementation.

[0310] The Boolean operation may be a NAND operation, or another Boolean operation.

[0311] The one or more Boolean operations may include a producer Boolean operation applicable to the producer hash value, and a consumer Boolean operation applicable to the consumer hash value.

[0312] Fig.24 An array of PEs is illustrated, and local input and output ports 5112 of PE 5102 are illustrated (one row per neighbor PE), two input ports 5111 of PE 5101 for receiving broadcasts, one output RC port 5114 of PE 5103, and two input RC ports 5113 of PE 5103. A single PE may have all three types of ports, but for ease of explanation, different types of ports are shown with respect to different PEs.

[0313] Fig.25An example of a PE 510 is illustrated, which may include Q1 local input ports 511(1)-511(Q1) and Q1 output ports 512(1)-512(Q1) for local communication with Q1 neighbors, and may include I / O ports for remote communication (RC) (e.g., with one or more Benes networks). The I / O ports for RC may include Q2 input ports 513(1)-513(Q2) and Q3 output ports 514(1)-514(Q3) (e.g., 1 unicast BN output port and 2 unicast BN input ports, but other numbers may be provided). The PE may include Q4 input ports 515(1)-515(Q4) for receiving broadcasts (provided simultaneously to a group of PEs, such as a row of PEs or a column of PEs in the array or any 2D group of PEs). PE 510 may also include input circuitry 521 (e.g., a plurality of multiplexers and control signals) for receiving content from an input port and transmitting the content to a computational core, which may include an arithmetic logic unit (ALU) 450 and registers 550(0)-550(15) of a register file 550. The content of the computational core and / or the content from the input circuitry may be fed to output circuitry 522 (e.g., a plurality of multiplexers and control signals), which outputs the content from any of the output ports of PE 510.

[0314] The subject matter which is regarded as the embodiments of the present disclosure is particularly pointed out and distinctly claimed in the concluding portion of the specification. However, the organization and method of operation of the embodiments of the present disclosure, together with objects, features, and advantages thereof, may be best understood by reference to the following detailed description when read in conjunction with the accompanying drawings.

[0315] It should be understood that for simplicity and clarity of illustration, the elements shown in the drawings are not necessarily drawn to scale. For example, for clarity, the size of some of the elements may be exaggerated relative to other elements. In addition, where deemed appropriate, reference numerals may be repeated in the drawings to indicate corresponding or similar elements.

[0316] Since most of the illustrated embodiments of the present disclosure can be implemented using electronic components and circuits known to those skilled in the art, the details will not be explained to a greater extent than necessary above in order to understand and appreciate the basic concepts of the embodiments of the present disclosure and so as not to confuse or distract from the teachings of the embodiments of the present disclosure.

[0317] Any reference to a method in this specification should be adapted to apply to a system capable of performing the method and should be adapted to apply to a computer-readable medium that is non-transitory and stores instructions for performing the method.

[0318] Any references in this specification to a system should be applied appropriately to a method executable by the system and should be applied appropriately to a computer-readable medium that is non-transitory and stores instructions executable by the system.

[0319] Any reference in this specification to a non-transitory computer-readable medium should be applied appropriately to methods applicable when executing instructions stored in the computer-readable medium, and should be applied appropriately to systems configured to execute instructions stored in the computer-readable medium.

[0320] The term "and / or" means additionally or alternatively.

[0321] The processing circuitry may be implemented as a central processing unit (CPU), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a full custom integrated circuit, a graphics processing unit (GPU), a hardware accelerator, a system on a chip, or the like.

[0322] References to any of the terms “comprise / comprises / comprising,” “including / includes,” and “may include” may be applied to any of the terms “consists / consisting” and “consisting essentially of.” For example, any of the methods describing steps may include more steps than shown in the figure, only the steps shown in the figure, or essentially only the steps shown in the figure. The same applies to components of a device, processor, or system and to instructions stored in any non-transitory computer-readable storage medium.

[0323] The subject matter may also be implemented in a computer program for running on a computer system, the computer program comprising at least code portions that, when run on a programmable device (such as a computer system), perform the steps of the method according to the subject matter or enable the programmable device to perform the functions of the apparatus or system according to the subject matter. The computer program may cause the storage system to assign the disk drives to the disk drive groups.

[0324] A computer program is a list of instructions such as a specific application or operating system. A computer program may include, for example, one or more of the following: a subroutine, a function, a procedure, an object method, an object implementation, an executable application, an applet, a service program, source code, object code, a shared library / dynamically loaded library, or other instruction sequence designed to be executed on a computer system.

[0325] The computer program may be stored internally on a non-transitory computer-readable medium. All or some of the computer programs may be provided on a computer-readable medium that is permanently, removably, or remotely coupled to an information processing system. The computer-readable medium may include, for example, but not limited to, any number of the following media: magnetic storage media, including magnetic disk and tape storage media; optical storage media such as optical disk media (e.g., CD-ROM, CD-R, etc.) and digital video disk storage media; non-volatile memory storage media, including semiconductor-based memory cells such as flash memory, EEPROM, EPROM, ROM; ferromagnetic digital memory; MRAM; volatile storage media, including registers, buffers or caches, main memory, RAM, etc.

[0326] A computer process typically consists of an executing (running) program or portion of a program, current program values ​​and state information, and resources used by an operating system to manage the execution of the process. An operating system (OS) is software that manages the sharing of a computer's resources and provides programmers with an interface for accessing those resources. An operating system processes system data and user input, and responds by allocating and managing tasks and internal system resources as a service to the system's users and programs.

[0327] A computer system may include, for example, at least one processing unit, associated memory, and a plurality of input / output (I / O) devices. When executing a computer program, the computer system processes information according to the computer program and generates resultant output information via the I / O devices.

[0328] In the foregoing specification, the subject matter has been described with reference to specific examples of embodiments of the subject matter. It will, however, be evident that various modifications and changes may be made therein without departing from the broader spirit and scope of the subject matter as set forth in the appended claims.

[0329] Furthermore, the terms "front," "rear," "top," "bottom," "above," "below," and the like, if any, in the specification and claims are used for descriptive purposes and not necessarily for describing permanent relative positions. It is to be understood that the terms so used are interchangeable under appropriate circumstances such that embodiments of the subject matter described herein are, for example, capable of operation in other orientations than those illustrated or otherwise described herein.

[0330] As discussed herein, connection can be any type of connection, which is suitable for, for example, transmitting signals from corresponding nodes, units or devices via intermediate devices or corresponding nodes, units or devices transmitting signals. Therefore, unless otherwise implied or indicated, connection can be, for example, direct connection or indirect connection. Connection can be illustrated or described with reference to a single connection, multiple connections, unidirectional connection or bidirectional connection. However, different embodiments can change the implementation of connection. For example, a separate unidirectional connection can be used instead of a bidirectional connection, and vice versa. Moreover, multiple connections can be replaced by a single connection that transmits multiple signals in serial or in a time multiplexing manner. Similarly, a single connection that carries multiple signals can be separated into various different connections that carry subgroups of these signals. Therefore, there are many options for transmitting signals.

[0331] Although specific conductivity types or polarities of potentials have been described in the examples, it should be understood that the conductivity types and polarities of potentials may be reversed.

[0332] Each of the signals described herein may be designed as either positive logic or negative logic. In the case of a negative logic signal, the signal is active low, where the logically true state corresponds to a logic level of zero. In the case of a positive logic signal, the signal is active high, where the logically true state corresponds to a logic level of one. Note that any of the signals described herein may be designed as either a negative logic signal or a positive logic signal. Thus, in alternative embodiments, those signals described as positive logic signals may be implemented as negative logic signals, and those signals described as negative logic signals may be implemented as positive logic signals.

[0333] Additionally, the terms "assert" or "set" and "negate" (or "deassert" or "clear") are used herein when referring to the splitting of a signal, status bit, or similar device into its logically true or logically false state. If the logically true state is a logic level one, the logically false state is a logic level zero. If the logically true state is a logic level zero, the logically false state is a logic level one.

[0334] Those skilled in the art will recognize that the boundaries between logic blocks are merely illustrative, and that alternative embodiments may merge logic blocks or circuit elements, or may impose alternative functional decompositions on various logic blocks or circuit elements. Therefore, it should be understood that the architectures depicted herein are merely exemplary, and that in fact many other architectures may be implemented that achieve the same functionality.

[0335] Any arrangement of components that achieve the same functionality is effectively "associated" so that the desired functionality is achieved. Thus, any two components combined herein to achieve a particular functionality may be considered "associated" with each other so that the desired functionality is achieved, regardless of architecture or intermediate components. Likewise, any two components so associated may also be considered "operably connected" or "operably coupled" to each other so as to achieve the desired functionality.

[0336] In addition, those skilled in the art will recognize that the boundaries between the above operations are merely illustrative. Multiple operations may be combined into a single operation, a single operation may be distributed in additional operations, and operations may be performed at least partially overlapping in time. In addition, alternative embodiments may include multiple instances of a particular operation, and the order of operations may be changed in various other embodiments.

[0337] The illustrated examples may be implemented as circuits located on a single integrated circuit or within the same device. Alternatively, the examples may be implemented as any number of separate integrated circuits or separate devices interconnected with each other in a suitable manner. The examples or portions thereof may be implemented as software or code representations of physical circuits or logical representations convertible to physical circuits, such as in any suitable type of hardware description language.

[0338] Example 1 is a method for processing the placement and routing of physical (PE) arrays, the method comprising: obtaining placement primitive (PP) location window information, the location window information defining possible locations of a set of PPs within a coarse-grained reconfigurable array (CGRA); obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be calculated by the PE array; and determining the location of the PE in the PE array based on the CG, the PP location window information and the hardware constraints.

[0339] In Example 2, the subject matter of Example 1 includes, wherein the PP position window information is determined before receiving the CG.

[0340] In Example 3, the subject matter of Examples 1-2 includes, wherein the PP location window information represents a local connectivity constraint.

[0341] In Example 4, the subject matter of Examples 1 to 3 includes, wherein the PP in the PP represents a consumer.

[0342] In Example 5, the subject matter of Examples 1 to 4 includes, wherein each PP of at least two PPs in the set of PPs represents connectivity between one or more producer PEs and one or more consumer PEs.

[0343] In Example 6, the subject matter of Examples 1 to 5 includes dividing the CG into sub-pictures.

[0344] In Example 7, the subject matter of Example 6 includes using integer linear programming to partition the CG into subgraphs.

[0345] In Example 8, the subject matter of Examples 6 to 7 includes partitioning the CG into subgraphs; and identifying long-range connectivity connections based on edges between subgraphs and based on the long-range connectivity constraints.

[0346] In Example 9, the subject matter of Example 8 includes, wherein the remote connectivity connection is a connection to a Benes network.

[0347] In Example 10, the subject matter of Examples 8-9 includes, wherein the splitting includes compensating for a delay introduced by the remote connectivity connection.

[0348] In Example 11, the subject matter of Examples 6 to 10 includes, wherein the local connectivity constraints define point-to-point connections between the PEs, wherein the point-to-point connections connect a PE to a neighbor of the PE.

[0349] In Example 12, the subject matter of Examples 6 to 11 includes fitting the subgraph to a bucket.

[0350] In Example 13, the subject matter of Example 12 includes laying out the PP within the window associated with the bucket by applying a plurality of layout iterations.

[0351] In Example 14, the subject matter of Example 13 includes, wherein the plurality of layout iterations are based on the PP position information.

[0352] In Example 15, the subject matter of Example 14 includes, wherein a window associated with a bucket does not exceed the bucket.

[0353] In Example 16, the subject matter of Examples 14-15 includes determining an order of the PPs to be evaluated during the plurality of layout iterations.

[0354] In Example 17, the subject matter of Example 16 includes, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

[0355] In Example 18, the subject matter of Examples 16-17 includes, wherein the plurality of layout iterations comprises a combination of checking positions of the PPs and arrangements of the PPs.

[0356] In Example 19, the subject matter of Examples 16 to 18 includes, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

[0357] In Example 20, the subject matter of Example 19 includes, wherein a set of layout iterations includes: (a) evaluating increasing combinations of PPs during different layout iterations of the set until a combination is found to be infeasible; (b) gradually reducing the combinations; and (c) gradually increasing the combinations.

[0358] In Example 21, the subject matter of Examples 19-20 includes, wherein a set of layout iterations comprises (a) evaluating increasing combinations of PPs during different layout iterations of the set until all subgraphs are laid out in the buckets.

[0359] In Example 22, the subject matter of Examples 19 to 21 includes, wherein a set of layout iterations includes: selecting a PP to be evaluated according to the order of the PPs to be evaluated to provide a selected PP; determining the feasibility of positioning the selected PP within the bucket and the position of the selected PP within the bucket is feasible, wherein the determination is based on the PP position information and based on a current state of the bucket, the current state of the bucket including nodes of the bucket that have been laid out in the bucket; and when there is a feasibility of positioning the selected PP within the bucket, selecting a new PP to be evaluated according to the order of the PPs to be evaluated to provide a new selected PP.

[0360] In Example 23, the subject matter of Examples 19 to 22 includes, wherein the determining is based on at least one of a local connectivity constraint or a remote connectivity constraint.

[0361] In Example 24, the subject matter of Examples 19 to 23 includes, wherein the set of layout iterations includes checking feasibility of layout of the combination of PPs based on the non-feasible layout information.

[0362] In Example 25, the subject matter of Example 24 includes, wherein checking the feasibility of the placement of the combination of PPs is further based on processing entity (PE) feasibility.

[0363] In Example 26, the subject matter of Example 25 includes, wherein the PE feasibility is based on a pre-layout of another PE.

[0364] In Example 27, the subject matter of Examples 16 to 26 includes, wherein the plurality of layout iterations comprises a combination of checking positions of the PPs.

[0365] In Example 28, the subject matter of Examples 16 to 27 includes, wherein the plurality of layout iterations includes examining combinations of permutations of the PPs.

[0366] In Example 29, the subject matter of Examples 14 to 28 includes, wherein a size of the window exceeds a corresponding size of the bucket.

[0367] In Example 30, the subject matter of Examples 12 to 29 includes, wherein the barrel includes at least two barrels that are different in shape from each other.

[0368] In Example 31, the subject matter of Examples 12 to 30 includes, wherein the bucket includes at least two buckets that are different in size from each other.

[0369] In Example 32, the subject matter of Examples 8 to 31 includes temporally balancing the CG before segmenting the CG into sub-pictures.

[0370] In Example 33, the subject matter of Example 32 includes, wherein the time balancing comprises using integer linear programming.

[0371] In Example 34, the subject matter of Examples 8 to 33 includes, wherein the splitting is performed while complying with the long-range connectivity constraint.

[0372] In Example 35, the subject matter of Example 34 includes, wherein the remote connectivity constraint limits the number of remote connections of a single PE.

[0373] In Example 36, the subject matter of Examples 34 to 35 includes, wherein the hardware constraint comprises different functions of at least two PEs in the PE array.

[0374] In Example 37, the subject matter of Examples 34 to 36 includes, wherein the set of PPs includes fully connected layout primitives.

[0375] In Example 38, the subject matter of Examples 1 to 37 includes forming the PE array by placing the PEs based on the determination of the positions of the PEs.

[0376] Example 39 is a method for processing order-based placement and routing of physical (PE) arrays, the method comprising: obtaining hardware constraints about the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing mathematical expressions to be calculated by the PE array; and performing a PP-based determination of the position of the PE in the PE array based on the computation graph, placement primitives (PP) and the hardware constraints; wherein the PP-based determination comprises determining an order of the PPs to be evaluated during the PP-based determination.

[0377] In Example 40, the subject matter of Example 39 includes obtaining PP position information about possible positions of a set of PPs within the window.

[0378] In Example 41, the subject matter of Example 40 includes, wherein each PP represents connectivity between one or more producer PEs and one or more consumer PEs.

[0379] In Example 42, the subject matter of Examples 40-41 includes, wherein at least one PP represents a producer.

[0380] In Example 43, the subject matter of Examples 40 to 42 includes segmenting the CG into sub-pictures.

[0381] In Example 44, the subject matter of Example 43 includes using integer linear programming to partition the CG into subgraphs.

[0382] In Example 45, the subject matter of Examples 43 to 44 includes partitioning the CG into subgraphs and defining edges between the subgraphs as long-range connectivity connections.

[0383] In Example 46, the subject matter of Example 45 includes, wherein the remote connectivity connection is a connection to a Benes network.

[0384] In Example 47, the subject matter of Examples 45 to 46 includes, wherein the splitting includes compensating for delay introduced by the remote connectivity connection.

[0385] In Example 48, the subject matter of Examples 43 to 47 includes, wherein the local connectivity constraints define point-to-point connections between the PEs, wherein the point-to-point connections connect a PE to a neighbor of the PE.

[0386] In Example 49, the subject matter of Examples 43 to 48 includes fitting the subgraph to a bucket.

[0387] In Example 50, the subject matter of Example 49 comprises laying out the PP within the window associated with the bucket by applying a plurality of layout iterations.

[0388] In Example 51, the subject matter of Example 50 includes, wherein the plurality of layout iterations are based on the PP position information.

[0389] In Example 52, the subject matter of Example 51 includes, wherein a window associated with a bucket does not exceed the bucket.

[0390] In Example 53, the subject matter of Examples 51 to 52 includes determining an order of the PPs to be evaluated during the plurality of layout iterations.

[0391] In Example 54, the subject matter of Example 53 includes, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

[0392] In Example 55, the subject matter of Examples 53 to 54 includes, wherein the plurality of layout iterations comprises a combination of checking positions of the PPs and arrangements of the PPs.

[0393] In Example 56, the subject matter of Examples 53 to 55 includes, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

[0394] In Example 57, the subject matter of Example 56 includes, wherein a set of layout iterations includes: (a) evaluating increasing combinations of PPs during different layout iterations of the set until a combination is found to be infeasible; (b) gradually reducing the combinations; and (c) gradually increasing the combinations.

[0395] In Example 58, the subject matter of Examples 56 to 57 includes, wherein a set of layout iterations comprises (a) evaluating increasing combinations of PPs during different layout iterations of the set until all subgraphs are laid out in the buckets.

[0396] In Example 59, the subject matter of Examples 56 to 58 includes, wherein a set of layout iterations includes: selecting a PP to be evaluated according to the order of the PPs to be evaluated to provide a selected PP; determining the feasibility of positioning the selected PP within the bucket and the position of the selected PP within the bucket is feasible, wherein the determination is based on the PP position information and based on a current state of the bucket, the current state of the bucket including nodes of the bucket that have been laid out in the bucket; and when there is a feasibility of positioning the selected PP within the bucket, selecting a new PP to be evaluated according to the order of the PPs to be evaluated to provide a new selected PP.

[0397] In Example 60, the subject matter of Examples 56 to 59 includes, wherein the set of layout iterations includes checking feasibility of layout of the combination of PPs based on the non-feasible layout information.

[0398] In Example 61, the subject matter of Examples 53 to 60 includes, wherein the plurality of layout iterations comprises a combination of checking positions of the PPs.

[0399] In Example 62, the subject matter of Examples 53 to 61 includes, wherein the plurality of layout iterations comprises checking a combination of permutations of the PPs.

[0400] In Example 63, the subject matter of Examples 51 to 62 includes, wherein a size of the window exceeds a corresponding size of the bucket.

[0401] In Example 64, the subject matter of Examples 49 to 63 includes, wherein the barrel includes at least two barrels that are different in shape from each other.

[0402] In Example 65, the subject matter of Examples 49 to 64 includes, wherein the bucket includes at least two buckets that are different in size from each other.

[0403] In Example 66, the subject matter of Examples 45 to 65 includes temporally balancing the CG before splitting the CG into sub-pictures.

[0404] In Example 67, the subject matter of Example 66 includes, wherein the time balancing includes using integer linear programming.

[0405] In Example 68, the subject matter of Examples 45 to 67 includes, wherein the splitting is performed while complying with the long-range connectivity constraint.

[0406] In Example 69, the subject matter of Example 68 includes, wherein the remote connectivity constraint limits the number of remote connections of a single PE.

[0407] In Example 70, the subject matter of Examples 68 to 69 includes, wherein the hardware constraint comprises different functionality of at least two PEs in the PE array.

[0408] In Example 71, the subject matter of Examples 68 to 70 includes, wherein the set of PPs includes fully connected layout primitives.

[0409] In Example 72, the subject matter of Example 71 includes forming the PE array by placing the PEs based on the determination of the positions of the PEs.

[0410] Example 73 is a method for processing a hash-based layout of a physical entity (PE) array, the method comprising: obtaining hardware constraints about the array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be calculated by the array; and performing a hash-based PP-based determination of the position of the PE in the array based on the computation graph, the layout primitive (PP) and the hardware constraints; wherein the hash-based PP determination comprises performing multiple hash-based layout iterations, wherein a set of hash-based layout iterations is associated with a portion of the array, wherein the set of hash-based layout iterations comprises: an initial hash-based layout iteration for initially filling an empty window with an initial allowed hardware implementation of an initial PP; and additional hash-based layout iterations for filling the partially filled window with one or more additional allowed hardware implementations of one or more additional PPs; wherein the additional hash-based layout iterations are performed based on one or more hash values ​​of one or more filled nodes.

[0411] In Example 74, the subject matter of Example 73 includes, wherein the additional hash-based layout iterations are further based on one or more hash values ​​of the additional enabled hardware implementations.

[0412] In Example 75, the subject matter of Example 74 includes, wherein the additional hash-based layout iteration comprises selecting an additional allowed hardware implementation from a set of additional allowed hardware implementations.

[0413] In Example 76, the subject matter of Example 75 includes, wherein the selecting is based on (i) one or more hash values ​​of members of the group and (ii) one or more hash values ​​of one or more populated nodes that at least partially appear in the members of the group.

[0414] In Example 77, the subject matter of Example 76 includes, wherein the one or more hash values ​​of the members of the group include a producer hash value and a consumer hash value.

[0415] In Example 78, the subject matter of Examples 76-77 includes, wherein the one or more hash values ​​of the members of the group include a producer hash value, a consumer hash value, and an idle hash value.

[0416] In Example 79, the subject matter of Examples 76 to 78 includes, wherein the one or more hash values ​​of members of the group include hash values ​​indicating producers, consumers, and idle nodes.

[0417] In Example 80, the subject matter of Examples 76 to 79 includes, wherein one or more hash values ​​of members of the group are canonical hash values.

[0418] In Example 81, the subject matter of Examples 76 to 80 includes, wherein one or more hash values ​​of members of the group are non-canonical hash values.

[0419] In Example 82, the subject matter of Examples 76 to 81 includes, wherein the selecting includes selecting a member of the group having one or more hash values ​​equal to the one or more hash values ​​of one or more populated nodes appearing in the member of the group.

[0420] In Example 83, the subject matter of Examples 76 to 82 includes, wherein the one or more hash values ​​of the members of the group are one or more bitmaps.

[0421] In Example 84, the subject matter of Examples 75 to 83 includes, wherein the selection of the additional allowed hardware implementations is followed by determining whether the additional allowed hardware implementations are placeable in the partially populated window.

[0422] In Example 85, the subject matter of Example 84 includes, wherein determining whether the additional enabled hardware implementation is placeable in the window comprises performing one or more Boolean operations.

[0423] In Example 86, the subject matter of Example 85 includes, wherein the Boolean operation of the one or more Boolean operations is applied to a hash value of the one or more populated nodes and to a hash value of the additionally enabled hardware implementation.

[0424] In Example 87, the subject matter of Example 86 includes, wherein the Boolean operation is a NAND operation.

[0425] In Example 88, the subject matter of Examples 85 to 87 includes, wherein the one or more Boolean operations include a producer Boolean operation applied to a producer hash value and a consumer hash value applied to a consumer hash value.

[0426] Example 89 is a method for processing the placement and routing of physical (PE) in an array of PEs, the method comprising: obtaining PP location information about possible locations of a set of placement primitives (PP) within a window; obtaining hardware constraints about the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be calculated by the PE array; wherein the PP location information is determined independently of the CG; and determining the position of the PE in the PE array based on the computation graph, the PP location information and the hardware constraints.

[0427] Example 90 is a non-transitory computer-readable medium for processing the placement and routing of physical (PE) devices in an array, the non-transitory computer-readable medium comprising instructions, the instructions being executed in response to a processor circuit of a computer-controlled device, causing the processor circuit to: obtain PP location information about possible locations of a set of placement primitives (PP) within a window; obtain hardware constraints about the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receive a computation graph (CG) representing a mathematical expression to be calculated by the PE array; wherein the PP location information is determined independently of the CG; and determine the position of the PE in the PE array based on the computation graph, the PP location information, and the hardware constraints.

[0428] Example 91 is a non-temporary computer-readable medium for processing the placement and routing of physical (PE) devices in an array, the non-temporary computer-readable medium comprising instructions, the instructions being executed in response to a processor circuit of a computer-controlled device, causing the processor circuit to: obtain hardware constraints regarding the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receive a computation graph (CG) representing mathematical expressions to be calculated by the PE array; and perform a PP-based determination of the position of the PE in the PE array based on the computation graph, a placement primitive (PP), and the hardware constraints; wherein the PP-based determination comprises determining an order of the PPs to be evaluated during the PP-based determination.

[0429] Example 92 is a non-transitory computer-readable medium for processing hash-based placement and routing of physical entities (PEs) in an array of PEs, the non-transitory computer-readable medium comprising instructions, the instructions being executed in response to a processor circuit of a computer-controlled device, causing the processor circuit to: obtain hardware constraints about the array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receive a computation graph (CG) representing a mathematical expression to be calculated by the array; and perform a hash-based PP-based determination of the location of the PE in the array based on the computation graph, the placement primitives (PP) and the hardware constraints; wherein the hash-based PP-based determination comprises performing multiple hash-based groups of placement iterations, wherein a hash-based group of placement iterations is associated with a portion of the array, wherein the hash-based group of placement iterations comprises: an initial hash-based placement iteration for initially filling an empty window with an initial allowed hardware implementation of an initial PP; and additional hash-based placement iterations for filling the partially filled window with one or more additional allowed hardware implementations of one or more additional PPs; wherein the additional hash-based placement iterations are performed based on one or more hash values ​​of one or more filled nodes.

[0430] Example 93 is at least one machine-readable medium comprising instructions that, when executed by a processing circuit, cause the processing circuit to operate to implement any of Examples 1 to 92.

[0431] Example 94 is an apparatus comprising means for implementing any of Examples 1-92.

[0432] Example 95 is a system for implementing any one of Examples 1 to 92.

[0433] Example 96 is a method for implementing any of Examples 1 to 92.

[0434] The present subject matter is not limited to physical devices or units implemented in non-programmable hardware, but is applicable to programmable devices or units capable of performing the desired device functions by operating according to appropriate program code, such as mainframes, minicomputers, servers, workstations, personal computers, notebooks, personal digital assistants, electronic games, automobiles and other embedded systems, cell phones and various other wireless devices, generally referred to in this application as "computer systems."

[0435] Other modifications, changes, and alternatives are also possible.Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense.

[0436] In the claims, any reference symbol placed in brackets shall not be interpreted as limiting the claim. The word "comprising" does not exclude the presence of other elements or steps than those listed in the claim. In addition, the term "a" or "an" as used herein is defined as one or more than one. Moreover, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be interpreted as implying that any particular claim containing such introduced claim elements is limited to the subject matter containing only one such element by introducing another claim element by the indefinite article "a" or "an", even if the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an". The same is true for the use of definite articles. Unless otherwise specified, terms such as "first" and "second" are used to arbitrarily distinguish the elements described by such terms. Therefore, these terms are not necessarily intended to indicate the time or other priority of such elements. The fact that certain measures are recited in mutually different claims does not mean that the combination of these measures cannot be used advantageously.

[0437] Although certain features of the subject matter have been illustrated and described herein, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is therefore to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true spirit of the subject matter.

Claims

1. A method for processing the placement and routing of PEs in a physical (PE) array, the method include: Obtaining placement primitive (PP) location window information, wherein the location window information defines possible locations of a set of PPs within a coarse-grained reconfigurable array (CGRA); Obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array; as well as The position of the PE in the PE array is determined based on the CG, the PP position window information and hardware constraints.

2. The method according to claim 1, wherein the PP position window information is determined before receiving the CG. The method of claim 1 , wherein the PP location window information represents a local connectivity constraint. The method according to claim 1 , wherein the PP in the PP stands for consumer. 5 . The method of claim 1 , wherein each PP of at least two PPs in the set of PPs represents connectivity between one or more producer PEs and one or more consumer PEs. The method according to claim 1 , comprising segmenting the CG into sub-graphs.

7. The method of claim 6, comprising using integer linear programming to partition the CG into sub-graphs.

8. The method according to claim 6, wherein include: Splitting the CG into sub-graphs; as well as Long-range connectivity connections are identified based on the edges between the subgraphs and based on the long-range connectivity constraints.

9. The method of claim 8, wherein the remote connectivity connection is a connection to a Benes network.

10. The method of claim 8, wherein the segmenting includes compensating for delay introduced by the remote connectivity connection.

11. The method of claim 6, wherein the local connectivity constraints define point-to-point connections between the PEs, wherein a point-to-point connection connects a PE to a neighbor of the PE.

12. The method of claim 6, comprising fitting the subgraphs into buckets.

13. The method of claim 12, comprising laying out the PP within a window associated with the bucket by applying a plurality of layout iterations. The method of claim 13 , wherein the plurality of layout iterations are based on the PP position information. The method of claim 14 , wherein a window associated with a bucket does not exceed the bucket.

16. The method of claim 14, comprising determining an order of PPs to be evaluated during the plurality of layout iterations. The method of claim 16 , wherein the plurality of layout iterations comprises a plurality of groups of layout iterations. The method of claim 16 , wherein the plurality of layout iterations comprises a combination of examining locations of PPs and arrangements of the PPs. The method of claim 16 , wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

20. The method of claim 19, wherein a set of layout iterations include: (a) evaluating increasing combinations of PPs during different placement iterations of the group until a combination is found to be infeasible; (b) gradually reducing the combinations; and (c) gradually increasing the combinations.

21. The method of claim 19, wherein a set of layout iterations comprises (a) evaluating increasing combinations of PPs during different layout iterations of the set until all subgraphs are laid out in the buckets.

22. The method of claim 19, wherein a set of layout iterations include: selecting the PP to be evaluated according to said order of the PP to be evaluated to provide a selected PP; determining the feasibility of locating the selected PP within the bucket and the location of the selected PP within the bucket is feasible, wherein the determination is based on the PP location information and based on a current state of the bucket, the current state of the bucket including nodes of the bucket that have been placed in the bucket; as well as When there is a feasibility of positioning the selected PP within the bucket, a new PP to be evaluated is selected according to the order of the PPs to be evaluated to provide a new selected PP.

23. The method of claim 19, wherein the determining is based on at least one of a local connectivity constraint or a remote connectivity constraint.

24. The method of claim 19, wherein a set of placement iterations includes checking feasibility of placement of a combination of PPs based on non-feasible placement information.

25. The method of claim 24, wherein checking the feasibility of the layout of the combination of PPs is further based on processing entity (PE) feasibility.

26. The method of claim 25, wherein the PE feasibility is based on a pre-layout of another PE.

27. The method of claim 16, wherein the plurality of layout iterations comprises examining a combination of locations of PPs.

28. The method of claim 16, wherein the plurality of layout iterations comprises examining combinations of permutations of PPs.

29. The method of claim 14, wherein a size of the window exceeds a corresponding size of the bucket.

30. The method of claim 12, wherein the barrel comprises at least two barrels that are different in shape from each other.

31. The method of claim 12, wherein the barrel comprises at least two barrels that are different in size from each other.

32. The method of claim 8, comprising temporally balancing the CG before segmenting the CG into sub-graphs.

33. The method of claim 32, wherein the time balancing comprises using integer linear programming.

34. The method of claim 8, wherein the segmentation is performed while complying with the long-range connectivity constraints.

35. The method of claim 34, wherein the remote connectivity constraint limits the number of remote connections of a single PE.

36. The method of claim 34, wherein the hardware constraints include different functions of at least two PEs in the PE array.

37. The method of claim 34, wherein the set of PPs comprises fully connected placement primitives.

38. The method of claim 1, comprising forming the PE array by placing the PEs based on the determination of the locations of the PEs.

39. A method for processing sequential-based placement and routing of PEs in an array of physical devices, the method include: Obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array; as well as A placement primitive (PP)-based determination of a location of the PE in the PE array is performed based on the computation graph, the PP, and the hardware constraints; wherein the PP-based determination includes determining an order of PPs to be evaluated during the PP-based determination.

40. The method of claim 39, comprising obtaining PP position information about possible positions of a set of PPs within a window.

41. The method of claim 40, wherein each PP represents connectivity between one or more producer PEs and one or more consumer PEs.

42. The method of claim 40, wherein at least one PP represents a producer.

43. The method of claim 40, comprising segmenting the CG into sub-graphs.

44. The method of claim 43, comprising using integer linear programming to partition the CG into sub-graphs.

45. The method of claim 43, comprising partitioning the CG into subgraphs and defining edges between subgraphs as long-range connectivity connections.

46. ​​The method of claim 45, wherein the remote connectivity connection is a connection to a Benes network.

47. The method of claim 45, wherein the segmenting comprises compensating for delay introduced by the remote connectivity connection.

48. The method of claim 43, wherein the local connectivity constraints define point-to-point connections between the PEs, wherein a point-to-point connection connects a PE to a neighbor of the PE.

49. The method of claim 43, comprising fitting the subgraphs into buckets.

50. The method of claim 49, comprising laying out the PP within a window associated with the bucket by applying a plurality of layout iterations.

51. The method of claim 50, wherein the plurality of layout iterations are based on the PP position information.

52. The method of claim 51 , wherein a window associated with a bucket does not exceed the bucket.

53. The method of claim 51, comprising determining an order of PPs to be evaluated during the plurality of layout iterations.

54. The method of claim 53, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

55. The method of claim 53, wherein the plurality of layout iterations comprises a combination of examining locations of PPs and arrangements of the PPs.

56. The method of claim 53, wherein the plurality of layout iterations comprises a plurality of groups of layout iterations.

57. The method of claim 56, wherein a set of layout iterations include: (a) evaluating increasing combinations of PPs during different placement iterations of the group until a combination is found to be infeasible; (b) gradually reducing the combinations; and (c) gradually increasing the combinations.

58. The method of claim 56, wherein a set of layout iterations comprises (a) evaluating increasing combinations of PPs during different layout iterations of the set until all subgraphs are laid out in the buckets.

59. The method of claim 56, wherein a set of layout iterations include: selecting the PP to be evaluated according to said order of the PP to be evaluated to provide a selected PP; determining the feasibility of locating the selected PP within the bucket and the location of the selected PP within the bucket is feasible, wherein the determination is based on the PP location information and based on a current state of the bucket, the current state of the bucket including nodes of the bucket that have been placed in the bucket; as well as When there is a feasibility of positioning the selected PP within the bucket, a new PP to be evaluated is selected according to the order of the PPs to be evaluated to provide a new selected PP.

60. The method of claim 56, wherein a set of placement iterations includes checking feasibility of placement of a combination of PPs based on non-feasible placement information.

61. The method of claim 53, wherein the plurality of layout iterations comprises examining a combination of locations of PPs.

62. The method of claim 53, wherein the plurality of layout iterations comprises examining combinations of permutations of PPs.

63. The method of claim 51 , wherein a size of the window exceeds a corresponding size of the barrel.

64. The method of claim 49, wherein the barrel comprises at least two barrels that are different in shape from each other.

65. The method of claim 49, wherein the barrel comprises at least two barrels that differ in size from each other.

66. The method of claim 45, comprising temporally balancing the CG before segmenting the CG into sub-graphs.

67. The method of claim 66, wherein the time balancing comprises using integer linear programming.

68. The method of claim 45, wherein the segmentation is performed while complying with the long-range connectivity constraints.

69. The method of claim 68, wherein the remote connectivity constraint limits the number of remote connections of a single PE.

70. The method of claim 68, wherein the hardware constraints include different functions of at least two PEs in the PE array.

71. The method of claim 68, wherein the set of PPs comprises fully connected placement primitives.

72. The method of claim 71, comprising forming the PE array by laying out the PEs based on the determination of the locations of the PEs.

73. A method for processing a hash-based layout of an array of physical entities (PEs), the method include: Obtaining hardware constraints on the array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the array; as well as performing a hash-based determination of a location of the PE in the array based on the computation graph, a placement primitive (PP), and the hardware constraints; wherein the hash-based determination of the PP comprises performing a plurality of hash-based sets of placement iterations, wherein a set of hash-based placement iterations is associated with a portion of the array, wherein the set of hash-based placement iterations comprises: an initial hash-based layout iteration for initial filling of empty windows with initial allowed hardware implementations of the initial PP; and additional hash-based placement iterations for filling the partially filled window with one or more additional enabled hardware implementations of one or more additional PPs; Wherein the additional hash-based layout iterations are performed based on one or more hash values ​​of the one or more populated nodes.

74. The method of claim 73, wherein the additional hash-based layout iterations are further based on one or more hash values ​​of the additional enabled hardware implementations.

75. The method of claim 74, wherein additional hash-based layout iterations include selecting an additional allowed hardware implementation from a set of additional allowed hardware implementations.

76. A method according to claim 75, wherein the selection is based on (i) one or more hash values ​​of members of the group and (ii) one or more hash values ​​of one or more populated nodes that appear at least partially in the members of the group.

77. The method of claim 76, wherein the one or more hash values ​​of members of the group include a producer hash value and a consumer hash value.

78. The method of claim 76, wherein the one or more hash values ​​of members of the group include a producer hash value, a consumer hash value, and an idle hash value.

79. The method of claim 76, wherein the one or more hash values ​​for members of the group include hash values ​​indicating producers, consumers, and idle nodes.

80. The method of claim 76, wherein one or more hash values ​​of members of the group are canonical hash values.

81. The method of claim 76, wherein one or more hash values ​​of members of the group are non-canonical hash values.

82. The method of claim 76, wherein the selecting comprises selecting a member of the group having one or more hash values ​​equal to the one or more hash values ​​of one or more populated nodes present in the member of the group.

83. The method of claim 76, wherein the one or more hash values ​​of the members of the group are one or more bitmaps.

84. The method of claim 75, wherein the selection of the additional allowed hardware implementations is followed by determining whether the additional allowed hardware implementations are placeable in the partially populated window.

85. The method of claim 84, wherein determining whether the additional allowed hardware implementations are placeable in the window comprises performing one or more Boolean operations.

86. The method of claim 85, wherein a Boolean operation of the one or more Boolean operations is applied to a hash value of the one or more populated nodes and to a hash value of the additional enabled hardware implementation.

87. The method of claim 86, wherein the Boolean operation is a NAND operation.

88. The method of claim 85, wherein the one or more Boolean operations include a producer Boolean operation applied to a producer hash value and a consumer hash value applied to a consumer hash value.

89. A method for processing the placement and routing of PEs in a physical (PE) array, the method include: Obtaining layout primitive (PP) position information about possible positions of a set of PPs within a window; Obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array; wherein the PP position information is determined independently of the CG; and A position of the PE in the PE array is determined based on the computation graph, the PP position information, and the hardware constraints.

90. A non-transitory computer readable medium for processing placement and routing of PEs in an array of PEs, the non-transitory computer readable medium comprising instructions that, in response to being executed by a processor circuit of a computer controlled device, cause the processor circuit to: Obtaining PP position information about possible positions of a set of layout primitives (PPs) within a window; Obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array; wherein the PP position information is determined independently of the CG; and The position of the PE in the PE array is determined based on the computation graph, the PP position information, and hardware constraints.

91. A non-transitory computer readable medium for processing placement and routing of PEs in an array of PEs, the non-transitory computer readable medium comprising instructions that, in response to being executed by a processor circuit of a computer controlled device, cause the processor circuit to: Obtaining hardware constraints on the PE array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the PE array; and A placement primitive (PP)-based determination of a location of the PE in the PE array is performed based on the computation graph, the PP, and the hardware constraints; wherein the PP-based determination includes determining an order of PPs to be evaluated during the PP-based determination.

92. A non-transitory computer readable medium for hash-based placement and routing of PEs in a processing entity (PE) array, the non-transitory computer readable medium comprising instructions that, in response to being executed by a processor circuit of a computer controlled device, cause the processor circuit to: Obtaining hardware constraints on the array, the hardware constraints comprising local connectivity constraints and remote connectivity constraints; receiving a computation graph (CG) representing a mathematical expression to be computed by the array ; as well as performing a hash-based placement primitive (PP) determination of a location of the PE in the array based on the computation graph, the PP, and the hardware constraints; wherein the hash-based determination of the PP comprises performing a plurality of hash-based sets of placement iterations, wherein a set of hash-based placement iterations is associated with a portion of the array, wherein the set of hash-based placement iterations comprises: An initial hash-based layout iteration for initial filling of empty windows with initial allowed hardware implementations of the initial PP; and Additional hash-based layout iterations for filling the partially filled windows with one or more additional enabled hardware implementations of one or more additional PPs; wherein the additional hash-based layout iterations are performed based on one or more hash values ​​of one or more filled nodes.