System and method for synthesizing network-on-chip for deadlock-free conversion
The clustering method for Network-on-Chip transforms networks into deadlock-free and near-optimal configurations, addressing cycle-induced deadlocks and optimizing resource usage, applicable to various network structures.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- ARTERIS INC
- Filing Date
- 2025-07-25
- Publication Date
- 2026-07-30
AI Technical Summary
The challenge in designing a Network-on-Chip (NoC) for System-on-Chip (SoC) is to create a connectivity map that adheres to physical constraints while avoiding cycles that can lead to deadlocks, while minimizing resource usage and optimizing wiring and switches.
A clustering method is applied to nodes and edges to transform the network into a deadlock-free and near-optimal configuration that maintains physical constraints, using edge and node clustering to reduce the number of wires and switches without introducing cycles.
The solution optimizes network resource usage and congestion, ensuring high throughput with efficient runtime performance and effective handling of incremental changes, applicable to irregular and regular networks like rings, meshes, and torii.
Smart Images

Figure 0007897995000001 
Figure 0007897995000002 
Figure 0007897995000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application is a continuation - in - part of U.S. Non - Provisional Patent Application No. 16 / 728,335, titled "PHYSICALLY AWARE TOPOLOGY SYNTHESIS OF A NETWORK", filed on December 27, 2019 by Moez CHERIF et al., the entire disclosure of which is incorporated herein by reference.
[0002] Field of the Invention This technology is in the field of computer system design, and more particularly, relates to topology synthesis for generating deadlock - free Network - on - Chip (NoC).
Background Art
[0003] Background Multiprocessor systems implemented within a System - on - Chip (SoC) communicate via a network such as a Network - on - Chip (NoC). Intellectual property (IP) blocks or elements or cores are used in chip design. The SoC includes instances of intellectual property (IP) blocks. Some IP blocks are masters. Some IP blocks are slaves. Masters and slaves communicate via a network such as a NoC.
[0004] Transactions in the form of packets are sent from a master to one or more slaves using any of a number of industry - standard protocols. A master connected to the NoC uses an address to select a slave and sends a request transaction to the slave. The NoC decodes the address and forwards the request from the master to the slave. The slave processes the transaction and sends a response transaction that is sent back to the master by the NoC.
[0005] The design of a Network of Controls (NoC) that handles all communication between all masters and their corresponding slaves involves establishing a connectivity mapping of the NoC within the floor plan. The challenge is that the connectivity map must take into account the location of IP blocks in the floor plan, which represents physical constraints in the floor plan. Furthermore, for NoCs, the connectivity map should avoid generating cycles. Cycles can lead to undesirable deadlock conditions where nodes along the cycle are in a cyclic "waiting" state, preventing them from accessing resources and sending messages to each other. Therefore, systems and methods are needed for the synthesis and transformation of networks. The process should minimize resource usage to generate a near-optimal, cycle-free network in light of physical constraints. The system and method should transform a given network into another functionally equivalent network with less wiring (e.g., fewer links) and fewer logical elements (e.g., fewer switches). Furthermore, the transformation must adhere to the connectivity constraints of the network and should not introduce new cycles that could lead to deadlocks. [Overview of the Initiative] [Means for solving the problem]
[0006] Summary of the Invention According to various embodiments and aspects of the present invention, while maintaining network connectivity constraints Furthermore, systems and methods for generating near-optimal networks, such as network-on-chip (NoC), are disclosed. According to various aspects and embodiments of the present invention, the system applies a clustering method to nodes and edges. Clustering transforms the network and generates a deadlock-free and near-optimal network that adheres to the physical constraints of the input network's floor plan and specifications.
[0007] One advantage of the present invention is to optimize the network and reduce resource usage and congestion. Another advantage is to use a deadlock-aware process to reduce the number of wires (edges) and switches (nodes) in the network. Another advantage is to produce optimal results when combined with the use of physical roadmap techniques. Another advantage is to generate a near-optimal or optimal network that maintains a cycle-free construction of the generated network while all transformations converge to a better-routed wiring result. Another advantage is the system's ability to apply this embodiment to any structure of irregular and regular networks, including rings, meshes, and torii. Another advantage is high throughput because the system implements the process with high runtime efficiency. Another advantage is its effectiveness in handling incremental changes during the synthesis process performed by the system.
[0008] Brief explanation of the drawing To better understand the present invention, refer to the accompanying drawings. The present invention is described in accordance with the aspects and embodiments of the following description with reference to the drawings or figures (FIG.), where similar reference numerals represent the same or similar elements. These drawings should not be considered as limiting the scope of the present invention, and the aspects and embodiments of the present invention described herein, as well as the best mode understood herein, are described in further detail by using the accompanying drawings. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows a network synthesis and transformation process for generating a new network, according to various aspects and embodiments of the present invention. [Figure 2] This figure shows a network including three source nodes and thirteen sink nodes, implemented in a floor plan with physical constraints, according to various aspects and embodiments of the present invention. [Figure 3]This figure shows a connectivity map including nodes and edges for the network shown in Figure 2, according to various aspects and embodiments of the present invention. [Figure 4] This figure shows a map of an input network containing three trunks, which is transformed into a new network represented by the new map using brute-force clustering. [Figure 5] This figure shows a map of an input network containing three trunks, which is transformed and synthesized to generate a new network represented by a new map, according to various aspects and embodiments of the present invention. [Figure 6] This figure shows a new network map, which is a map, that is transformed to generate another new network represented by the new map, according to various aspects and embodiments of the present invention. [Figure 7] This figure shows a map, which is a new network, that is transformed to generate another new network represented by the new map, according to various aspects and embodiments of the present invention. [Figure 8] Figure 7 shows a map of a new network according to various aspects and embodiments of the present invention, in which edge clustering is not implemented to avoid cycles. [Figure 9] This figure shows a map, which is the network in Figure 8, that is transformed to generate another new network represented by a new map, according to various aspects and embodiments of the present invention. [Figure 10A] This figure shows a flow process for performing edge cluster grouping according to various aspects and embodiments of the present invention. [Figure 10B] This figure shows a flow process for performing node cluster grouping according to various aspects and embodiments of the present invention. [Figure 11] This figure shows a map of edge clustered networks having possible node clusters according to various aspects and embodiments of the present invention. [Figure 12]This figure shows the steps for collapsing a possible node cluster according to various aspects and embodiments of the present invention. [Figure 13] This figure shows an input network having node clusters that may be collapsible, as well as node clusters that may not be collapsible due to network constraints, according to various aspects and embodiments of the present invention. [Figure 14] Figure 13 shows the input network after a possible node cluster has been collapsed, according to various aspects and embodiments of the present invention. [Modes for carrying out the invention]
[0010] Detailed explanation The following describes various embodiments of the present invention, illustrating various aspects and embodiments of the present invention. In general, the embodiments can be used in any combination of the described aspects. All descriptions herein, listing principles, aspects, and embodiments, as well as specific examples thereof, are intended to encompass both their structural and functional equivalents. Furthermore, such equivalents are intended to include both currently known equivalents and equivalents to be developed in the future, i.e., any elements being developed that perform the same function regardless of their structure.
[0011] Where used herein, the singular forms “a,” “an,” and “the” refer to multiple objects unless the context otherwise explicitly indicates otherwise. Throughout this specification, any reference to “one aspect,” “one aspect,” “a particular aspect,” “various aspects,” or similar phrases means that a particular aspect, feature, structure, or characteristic described in relation to any embodiment is included in at least one embodiment of the present invention.
[0012] Throughout this specification, the appearances of the phrases "in one embodiment", "in at least one embodiment", "in an embodiment", "in certain embodiments", and similar language, while not necessarily all referring to the same embodiment, can refer to the same or similar embodiments. Further, the aspects and embodiments of the invention described herein are merely illustrative and should not be construed as limiting the scope or spirit of the invention as understood by those skilled in the art. The disclosed invention is effectively made or used in any embodiment that includes any novel aspect described herein. All descriptions in this specification enumerating aspects and embodiments of the invention are intended to encompass both their structural and functional equivalents. Such equivalents are intended to include both currently known equivalents and equivalents developed in the future.
[0013] As used herein, "master" and "initiator" refer to similar intellectual property (IP) blocks, units, or modules. The terms "master" and "initiator" are used interchangeably within the scope and embodiments of the invention. As used herein, "slave" and "target" refer to similar IP blocks. The terms "slave" and "master" are used interchangeably within the scope and embodiments of the invention. As used herein, a transaction can be a request transaction or a response transaction. Examples of request transactions , include write requests and read requests.
[0014] As used herein, a node is defined as a distribution point or communication endpoint that can create, receive, and / or transmit information via a communication path or channel. A node can refer to any one of a switch, splitter, merger, buffer, and adapter. As used herein, a splitter and a merger are switches. However, not all switches are splitters or mergers. As used herein, according to various aspects and embodiments of the present invention, the term "splitter" represents a switch having a single input port and a plurality of output ports. As used herein, according to various aspects and embodiments of the present invention, the term "merger" represents a switch having a single output port and a plurality of input ports.
[0015] According to various aspects and one embodiment of the present invention, synthesis and transformation are performed on a deadlock-free network as described herein. The resulting transformed network topology is also cycle-free. As used herein, a "cycle-free" network is a network that has no route or path that traverses the same node twice. The terms "path" and "route" are used interchangeably herein. A path includes any combination of nodes and edges (also referred to as links herein), is composed of them, and along which data moves from a source to a destination. According to various aspects and embodiments of the present invention, the following notations are defined.
[0016] E is any edge or link. LE is the longest edge.
[0017] BE is a set of reserved edges. NBE is a set of non-reserved edges.
[0018] CL is an edge cluster or a cluster of links. G(CL) is the total gain (or cost) of CL.
[0019] WL is the wiring length of an edge or link. WL(CL) is the total wiring length of the cluster CL.
[0020] Referring here to Figure 1, a process 100 according to various embodiments of the present invention is shown for synthesizing and transforming networks to generate a near-optimal or optimal cycle-free network. According to one embodiment of the present invention, the input network is cycle-free. According to various other embodiments of the present invention, the input network is any network, which may have cycles, irregular topologies, and regular topologies (e.g., mesh, ring, torii, etc.). Network.
[0021] The synthesis / transformation process includes edge clustering and node clustering. The synthesis process minimizes resource usage and generates a near-optimal cycle-free network. The resulting network structure is subject to the physical constraints of the floor plan described as part of this specification. According to some aspects of the present invention, the synthesis process includes optimizing the objective function and optimizing the global cost corresponding to the total wiring length (representing links) at the edges of the network. According to various aspects and embodiments of the present invention, the two clustering phases (edge clustering and node clustering) are performed on the input network. The synthesis process operates while maintaining the network cycle-free, and converges to an optimal structure. This is achieved using two network transformations: edge clustering and node clustering. These transformations take an existing network as input and produce a network that is more optimal than the input network according to some metric as output. These transformations do not introduce cycles into the newly generated network, which is a significant outcome and beneficial because cycles cause deadlocks as described. According to various aspects and embodiments of the present invention, the synthesis process reconfigures the network to eliminate cycles if present, and then applies clustering and optimization while adhering to and maintaining cycle-free properties. Furthermore, the physical constraints of the floor plan are maintained and adhered to.
[0022] As shown in Figure 1, the synthesis process 100 shows an overall synthesis flow including two transformation or clustering phases. The input network 102 is arbitrary. The input network 102 can be designed manually (e.g., by a human) or can be manufactured (automatically designed) by a computer-aided design tool. According to one aspect and embodiment of the present invention, the input network 102 is used, and the synthesis process 100 starts with a cycle-free network topology 110. As described above, the network topology 110 can be any network according to several other aspects and embodiments of the present invention. The synthesis process 100 uses a transformation module 112 to perform an edge clustering transformation on the network topology 110 in order to generate an edge clustered network topology. The synthesis process 100 uses an optimization module 114 to optimize the edge clustered network topology. The synthesis process 100 uses a transformation module 116 to perform node clustering on the optimized network topology. The resulting network topology 120 is a cycle-free network topology. According to one aspect of the present invention, the resulting network topology 120 may be processed by the user for further optimization and design finishing used to generate the final network topology 130. According to another aspect of the present invention, the resulting network topology 120 may be processed by a synthesis tool for further optimization and design finishing used to generate the final network topology 130. As described above, according to some other aspects and embodiments of the present invention, the resulting network topology 120 is a cycle-free network topology, even if the initial network topology, such as network topology 102, had cycles.
[0023] Referring here to Figure 2, a floor plan 200 is shown according to various aspects and embodiments of the present invention, including nine restricted areas 210 and three source nodes 220 communicating with thirteen sink nodes 230. The floor plan 200 shows direct communication connections using direct edges, such as edges (or links) 240 between the source nodes 220 and the sink nodes 230. As shown, restricted areas 210 represent space on the floor plan that links or edges cannot traverse, and links must be located in areas on the floor plan not occupied by restricted areas 210. According to various aspects and embodiments of the present invention, the system receives as input any network structure that implements a floor plan 200 having connectivity between the source nodes 230 and the sink nodes 240. The floor plan 200 is provided to the system as a map. The system processes the map using edge and node clustering to generate a more optimized structure.
[0024] According to various aspects and embodiments of the present invention, the input network has cycles. Therefore, edge and node clustering is not intended to disrupt existing cycles. The system optimizes the network without increasing the number of existing cycles. According to various aspects and embodiments of the present invention, the input network is cycle-free. Therefore, the optimized network is also cycle-free.
[0025] Referring here to Figure 3, an example of a connected network 300 implementing the communication connectivity (or links) of the floor plan 200 in Figure 2 is shown. Nodes are switches and other active network elements. According to various aspects and embodiments of the present invention, the connected network 300 includes switches such as switch 350. According to various aspects and embodiments of the present invention, the connected network 300 includes edges such as edge 360. Edges are links between nodes, and links are bundles of electrical connections. Nodes and links constitute a path. According to various aspects and embodiments of the present invention, many nodes are in close proximity to each other. According to various aspects and embodiments of the present invention, many edges have similar shapes so that edges can begin and end at adjacent nodes and their profiles can be assimilated. The system combines nodes and edges that can be assimilated. According to various aspects and embodiments of the present invention, assimilated edges and nodes are combined to obtain an improved network in order to reduce the number of network wiring and logical elements while maintaining equivalent functionality.
[0026] Referring here to Figure 4, pre-clustering map 400 and post-clustering map 450 are shown according to various aspects and embodiments of the present invention. Maps 400 and 450 include switches and links. For illustrative purposes, only a portion of the network 300 in Figure 3, which is the input network and the structure on which clustering operates, is shown. According to various aspects and embodiments of the present invention, map 400 is shown having multiple switches, such as switch 402, and multiple links, such as link 404. Map 450 of a subnetwork shows the clustering result when applied to three trunks 410, 412, and 414. According to various aspects and embodiments of the present invention, map 450 is shown having multiple switches, such as switch 452, and multiple links, such as link 454. According to various aspects and embodiments of the present invention, switch 402 can be assimilated. Assimilation of switch 402 results in switch 452. This is the process of node clustering. According to various aspects and embodiments of the present invention, link 404 can be assimilated. Assimilation of link 404 results in link 454. This process is referred to as edge clustering.
[0027] Map 400 represents a subnetwork containing three separate trunks 410, 412, and 414 of an input network such as network 300. Nodes and edges may be stacked on top of each other or spaced apart as shown in the diagram. According to various aspects and embodiments of the present invention, each trunk 410, 412, and 414 does not have a cycle, since the input network such as network 300 is cycle-free. As described herein, according to various aspects and embodiments of the present invention, the edge clustering and node clustering processes operate similarly for a network that has a cycle, which may be the input network.
[0028] According to various aspects and embodiments of the present invention, clustering generates a more compact and optimized structure with respect to resource usage (cable length, performance, etc.) and keeps the network cycle-free. This is achieved by clustering and folding “similar” edge and neighboring nodes. Network map 450 is cycle-free and optimal and implements the same local connectivity map as map 400, where the local connectivity map is a map between inbound and outbound points to the trunk. Clustering maintains connectivity locally and globally (i.e., between source and sink) throughout all transformations. For example, switch 406 can be clustered to obtain switch 456. Trunk 414 is clustered by brute force When applied, it appears to have a loop shape that can generate cycles, and therefore deadlocks. According to various aspects and embodiments of the present invention, clustering is applied only when no cycles are introduced, in order to prevent this possibility of cycles and deadlocks. To maximize the benefits, this process mainly reduces long edges or links.
[0029] Referring here to Figure 5, a map 500 of a pre-cluster network having edge clusters or cluster link (CL) groupings such as CL1, CL2, CL3, CL4, and CL5, and three trunks 510, 512, and 514, according to various aspects and embodiments of the present invention. The CL groups are possible groups for performing edge clustering, each of which can be implemented in any order according to the specified purposes of various aspects and embodiments of the present invention. A map 550 of a post-cluster network in which clustering has been performed on the edges within the CL1 grouping is shown.
[0030] According to various aspects and embodiments of the present invention, one objective and focus of performing edge clustering is to minimize long edges. Many long edges traversing narrow passages between two or more restricted areas can lead to wiring congestion. Minimizing the wiring of long edges contributes to reducing congestion. According to various aspects and embodiments of the present invention, the length of an edge (link) is measured as the length of the wiring between the endpoints of the edge.
[0031] According to various aspects and embodiments of the present invention, all edges (links) are initially marked as unreserved. An edge (or link) is considered "reserved" if it has already been selected and assigned to a cluster or CL of edges. For example, link 504 is a reserved link because it has been selected and assigned to CL1.
[0032] According to one embodiment of the present invention, edge clustering operates iteratively, with each iteration applying two main steps: (1) building an edge cluster such as CL1, and (2) collapsing the edges (links of CL1) to implement the cluster.
[0033] Figure 5 shows a map 500 of the identified cluster set. The process outlined below according to various aspects and embodiments of the present invention identifies all possible edge clusters labeled as CL1, CL2, CL3, CL4, and CL5 in map 500. A possible cluster is a group of edges that are close to each other and point in the same direction. For clarity and simplicity, an exemplary embodiment of cluster CL1 is described below, and the process outline is given below. According to various aspects and embodiments of the present invention, map 550 shows the result of collapsing CL1, with three links 504 collapsing to result in link 562 and two new nodes 564 and 566. More specifically, the embodiment of cluster CL1 includes removing all edges 504 of CL1 and inserting a single edge 562 connected to nodes 564 and 566. To maintain network connectivity of the input network, all starting points of the original edges of CL1 are connected to node 564, and all endpoints of the edges of CL1 are connected from node 566. Therefore, instead of using the three long edges 504 of CL1, after the implementation of cluster grouping CL1, the network uses six small edges and one long edge or link 562 to maintain the original connectivity at this point. In this simple example, the cost of the cluster was the density of edges. According to one embodiment of the present invention, the process is to begin implementation with the largest cluster first, and so on. According to various aspects and embodiments of the present invention, if the cost is the gain with respect to the wiring length, the cluster The implementation of the CL3 grouping began with CL3 due to the wiring length of each edge in the CL3 grouping.
[0034] As described according to various aspects of the present invention, the edge clustering process involves iteratively grouping edges into individual clusters such as CL1, CL2, CL3, CL4, and CL5. Once the grouping of edges or CLs is identified, each CL is ranked with respect to gain, which is the gain in terms of how much reduction in trace length is achieved and / or how much performance is improved. The list of clusters (CLs) is then sorted in descending order of the calculated gains. The sorted list is then traversed to select the best cluster for implementation. Once this is done, there are two possible cases. According to various aspects and embodiments of the present invention, if all edges in a cluster group are found to be acceptable and suitable, the implementation is valid and there is no need to update the remaining clusters. The process selects the next best cluster in the sorted list and proceeds with its implementation.
[0035] According to various aspects and embodiments of the present invention, if there is an edge that is rejected for introducing a new cycle or for breaking some of the specified constraints, the process removes that edge (link) from the current cluster. The process identifies whether the removed edge can be grouped in the next cluster. This ensures that all edges are considered for clustering and optimization.
[0036] The implementation of a cluster of suitable, acceptable edges operates by considering all edges of the network or subnetwork provided as input. The process traverses all edges. For each edge, the process identifies whether folding the edge by other cluster edges could result in a cycle. Cycles are identified by a graph search across the network, finding paths connecting the preceding and succeeding points of the edge's endpoints. Edges that introduce cycles are removed or excluded from the cluster. According to various aspects and embodiments of the invention, once a cluster is fully validated as cycle-free, the cluster is implemented. The process then continues building and implementing the next cluster, and so on, until all edges have been considered. According to one aspect and embodiment of the invention, to maintain runtime efficiency, the process checks for cycle-free acceptance only when a cluster has been picked up for implementation. The advantage achieved is to avoid disqualifying good edges from cluster grouping with others early in the process.
[0037] The steps for building a cluster function iteratively. In each iteration, the process first creates an empty cluster CL and selects the longest edge (LE) from a set of unreserved edges (NBEs). Next, the process traverses the NBE, extracting edges moving in the same direction as the LE, with endpoints near the LE's endpoints. Building a cluster around the LE is an iterative process based on recalculating the centroid of the edges and recognizing / assimilating new edges in their vicinity. This approach has the advantage of better covering common cases of non-vertical and non-horizontal edges.
[0038] According to various aspects and embodiments of the present invention, only edges that are adjacent in the same direction and do not introduce cycles are retained in the cluster group. Once all edges are marked as reserved and moved from the NBE set to the BE set, the cluster is ready to be implemented and the edges can be folded.
[0039] Referring now to Figures 6, 7, and 9, edge clustering for the remaining cluster groups CL2, CL3, and CL5 according to various aspects and embodiments of the present invention. The process for performing the clustering is shown. In Figure 6, edge clustering is performed on CL2, resulting in edge 610 and nodes 620 and 630. In Figure 7, edge clustering is performed on CL3, resulting in edge 710 and nodes 720 and 730. In Figure 9, edge clustering is performed on CL5, resulting in edge 910, node 920, and node 930.
[0040] Referring to Figure 8, edge cluster CL4 includes edges 810 and 820. Implementing an edge cluster for CL4 would result in only one edge instead of edges 810 and 820; therefore, performing cluster grouping for CL4 would create a cycle caused by edge 820. Consequently, no implementation is performed for CL4, the paths are isolated, and the network's cycle-free properties are maintained. According to various aspects of the present invention, if there are other edges in the region of edge 810 and edge 810 can be selected to be part of another cluster, then edge 810 is selected as part of another cluster and collapsed. This process is dynamic and continues to converge toward an optimal solution as long as there are alternative cluster grouping options and room in the floor plan.
[0041] Referring here to Figure 10A, a process for generating edge cluster groupings such as CLs according to various embodiments of the present invention is shown. In step 1010, the system receives as input a network (or subnetwork) that is cycle-free and assigns all edges to a set of NBEs. In step 1012, the system generates a new empty cluster group CL to add edges to. NewThe system creates the CL. In step 1014, the system traverses the network edges, selects an LE from the set of NBEs, and the selected LE is added to the CL. In step 1016, the system identifies all NBEs close to the selected LE and adds the adjacent edges to the CL. In step 1018, the system removes any adjacent edges added in step 1016 from the CL if the adjacent edges introduce a cycle or violate any network constraints. As a result, an updated CL is obtained, and the set of edges can be implemented during the cluster implementation step. In step 1020, the system moves the edges that are part of the CL from the set of NBEs to the set of BEs. In step 1022, the system performs cluster grouping and replaces the edges in the CL with two new nodes N1 and N2 and the new edges. Nodes N1 and N2 are connected by the new edges. In step 1024, the system uses short edges to connect all nodes at the beginning of the CL to node N1. In step 1026, the system connects all nodes at the end of the CL to node N2. In step 1028, the system determines whether the NBE has other edges or whether the set of NBEs is empty. If the set of NBEs is not empty, the system repeats the cluster grouping process by returning to step 1012. If the set of NBEs is empty, there are no other edges in the set of NBEs, and the process terminates.
[0042] Referring here to Figure 10B, a process for generating node clusters according to various embodiments of the present invention is shown. In step 1070, the system receives a network with nodes (and links) and constraints as input. The input network is an edge clustered network, as outlined below according to various embodiments of the present invention. According to various embodiments of the present invention, the input network is any network including a network that has not been transformed using edge clustering. In step 1072, according to various embodiments of the present invention, the nodes are traversed to determine or identify at least two nodes that can be combined to form a possible node cluster. According to one embodiment of the present invention, in step 1072, all nodes are traversed. According to one embodiment of the present invention, step In step 1072, the nodes are traversed until at least two nodes that can be combined to form a possible node cluster are identified. In step 1074, the identified nodes are selected to form a possible node cluster. In step 1076, according to some aspects of the present invention, the remaining nodes are traversed. In step 1078, if other nodes that can be combined with the possible node cluster are identified, in step 1080 the identified nodes are added to the possible node cluster. In step 1078, if no other nodes are identified to add to the possible node cluster, in step 1082 the possible node cluster is folded as outlined below to form a new folded node. In step 1084 the system uses the new folded node to generate a transformed network. According to various aspects of the present invention, the system repeats the process for the transformed network until all different possible node clusters are identified. Each iteration of the process may result in a new transformed network, and the process may be repeated for the new transformed network resulting from the previous iteration, which is outlined in detail below according to various aspects and embodiments of the present invention disclosed herein.
[0043] Referring here to Figure 11, an edge clustered network 1100 according to various aspects and embodiments of the present invention is shown. The network 1100 includes three trunks and eight possible node clusters, such as node cluster 1110. The system performs a transformation referred to as node clustering. A node cluster is a group of nodes that are fitted together and located together. A fitted node pair is a pair of nodes that does not introduce a new cycle into the network, so that the resulting nodes do not exceed the maximum boundary of the number of I / O ports of the nodes, which is a parameter of this method. Fitted pairs should also adhere to performance targets, if the system takes performance constraints and metrics into consideration.
[0044] Matching nodes, such as switch elements, are grouped (clustered) to create a network using fewer resources, such as fewer logic elements and fewer wires. The node clustering process operates in an iterative and multipath manner. The system traverses a list of nodes. The system groups nodes into possible clusters based on their proximity in the floor plan. The system uses an iterative process, starting with one node and continuing to add new nodes to possible clusters, taking into account the "Manhattan ball" around the bucket centroid. When the system can no longer add new nodes to a possible cluster, the possible cluster is considered fully formed. The system then proceeds to start with a new node that is not yet in any of the previously built possible clusters.
[0045] Using possible node clusters, the system traverses the possible node clusters and performs a cost analysis on each with respect to a score function. According to one aspect of the present invention, the score function is based on the cluster size. According to one aspect of the present invention, once all possible node clusters have been analyzed and cost values have been obtained, the system sorts the possible node clusters in descending order of their cost. According to various other aspects of the present invention, the possible node clusters can be sorted in other ways, and the scope of the present invention is not limited thereto.
[0046] The system traverses a sorted list of possible node clusters, processing them one at a time. For the currently selected possible node cluster, the system iteratively identifies all matching pairs of nodes and scores them in terms of the gain they would yield if they were folded together. Various embodiments of the present invention According to this invention, cost is expressed in terms of WL. According to various aspects of this invention, cost is expressed in terms of performance. According to various aspects of this invention, cost is expressed in terms of the increase in merging nodes. According to various aspects of this invention, cost is expressed in terms of any combination of WL, performance, or increase. Once all pairs have been cost-calculated, the system selects the most suitable pair, removes the suitable pair from the list of candidate nodes, and performs folding of the suitable node pair.
[0047] According to various aspects of the present invention, folding two nodes N1 and N2 involves removing nodes N1 and N2. The removed nodes are then replaced with a new node N3. The system connects all preceding elements of N1 and N2 to N3. The system then connects N3 to all following elements of N1 and N2. The system updates all routes that passed through N1 and N2 with N3 in order to continue incrementally updating the routes. Once the system has updated all routes, the system then updates the list of candidate nodes with the new node N3. The system also updates the cost of the affected candidate nodes. The system then selects a new top-tier candidate pair and proceeds along the same scheme. The system iteratively repeats this process until all nodes are folded or there are no more acceptable pairs available for folding.
[0048] Referring now to Figure 12, the possible node cluster 1110 in Figure 11 is implemented according to various aspects and embodiments of the present invention to generate the resulting node 1210. The system selects two nodes that are a matching node pair from the possible node cluster 1110, as outlined above, and folds them to generate a new node. This results in a new possible node cluster 1202. The system repeats the implementation process for node cluster 1202 to generate another new possible node cluster 1204. This process is carried out for node cluster 1204 to generate node cluster 1206. The matching node pair in node cluster 1206 is implemented to generate node 1210. Once node 1210 is generated, there are no other nodes in the possible node cluster. According to various aspects of the present invention, if no further folding can be done within the current possible node cluster, the system moves on to the next possible node cluster. The system proceeds using the same method until all possible node clusters have been processed.
[0049] Referring here to Figure 13, input networks having possible node clusters 1310, 1320, 1330, 1340, 1350, 1360, 1370, and 1380 according to various aspects and embodiments of the present invention. The numbering of the possible node clusters in this example is arbitrary and does not indicate their ranking in a sorted list. All nodes in possible node clusters 1310, 1320, 1330, 1340, 1350, and 1360 are acceptable combinations that can be collapsed together. All nodes in possible node clusters 1370 and 1380 are not acceptable combinations because collapsing these possible node clusters would result in an unacceptable cycle.
[0050] Referring now to Figure 14, the final node clustering map 1400 for the input network in Figure 13 is shown. Map 1400 includes the collapsed clusters 1410, 1420, 1430, 1440, 1450, and 1460, as well as the nodes that were part of the uncollapsed node clusters 1370 and 1380 (in Figure 13). According to various aspects of the present invention, the resulting map 1400 is obtained by the system repeating the implementation step as many times as clustering can collapse existing nodes and create new nodes. The system concludes that, as is evident from map 1400, convergence can no longer be obtained and clustering can no longer be applied. Node clustering is stopped when no more nodes can be obtained. According to various aspects of the present invention, the multipathing performed by the system ensures that newly created nodes resulting from folding matching pairs of nodes are also considered for further expansion and folding with other nodes.
[0051] According to some aspects and embodiments, the tool can be used to ensure that multiple iterations of synthesis are performed for incremental optimization of the NoC. After the system implements and executes the synthesis process, the results are generated in a machine-readable format, such as a computer file, using a clearly defined format for capturing information. The scope of the present invention is not limited by any particular format.
[0052] Certain methods according to various aspects of the present invention can be performed by instructions stored in a non-temporary computer-readable medium. The non-temporary computer-readable medium stores code, which, when executed by one or more processors, causes a system or computer to perform steps of the methods described herein. The non-temporary computer-readable medium includes rotating magnetic disks, rotating optical disks, flash random access memory (RAM) chips, and other mechanically moving or solid-state storage media. Any type of computer-readable medium is suitable for storing code, which includes instructions according to various examples.
[0053] While specific examples have been described herein, it should be noted that different combinations of different components from different embodiments may be possible. Notable features are presented to better illustrate the embodiments. However, it is clear that certain features can be added, modified, and / or omitted without altering the functional aspects of these embodiments as described.
[0054] Various embodiments are methods that utilize the behavior of any or a combination of machines. Embodiments of a method are completed whenever most of the configuration steps are performed. For example, according to various aspects and embodiments of the present invention, an IP element or unit includes a processor (e.g., CPU or GPU), random access memory (RAM - e.g., off-chip dynamic RAM or DRAM), and network interfaces for wired or wireless connections such as Ethernet®, WiFi, 3G, 4G Long-Term Evolution (LTE), 5G, and other wireless interface standard radios. The IP may also include, as needed, various I / O interface devices such as keyboards and mice, in particular, for various peripheral devices such as touchscreen sensors, geolocation receivers, microphones, speakers, Bluetooth® peripherals, and USB devices. The processor performs the steps of the method described herein by executing instructions stored in the RAM device.
[0055] Some embodiments are one or more non-temporary computer-readable media configured to store instructions for the methods described herein. Any machine holding a non-temporary computer-readable media containing any of the required codes can carry out the embodiments. Some embodiments may be implemented as physical devices such as semiconductor chips, hardware description language representations of the logical or functional behavior of such devices, and one or more non-temporary computer-readable media configured to store such hardware description language representations. The descriptions herein enumerating principles, aspects, and embodiments encompass both their structural and functional equivalents. Elements described herein as being coupled have valid relationships that can be realized by direct connections or indirectly using one or more other intervening elements.
[0056] Those skilled in the art will recognize numerous modifications and variations. These modifications and variations include any relevant combination of the disclosed features. The descriptions herein enumerating principles, aspects, and embodiments encompass both their structural and functional equivalents. Elements described herein as “coupled” or “communicatively coupled” have an effective relationship that can be realized by direct connection or indirect connection using one or more other intervening elements. Embodiments described herein as “communicating” or “in communication with” another device, module, or element include any form of communication or link and include an effective relationship. For example, a communication link may be established using a wired connection, a wireless protocol, a near-field protocol, or RFID.
[0057] To the extent that the terms “including,” “includes,” “having,” “has,” and “with,” or variations thereof, are used in either the detailed description or the claims, such terms are intended to be as comprehensive as the term “comprising.”
[0058] Therefore, the scope of the present invention is not intended to be limited to the exemplary embodiments and aspects shown and described herein. Rather, the scope and spirit of the present invention are embodied in the appended claims.
Claims
1. A tool for optimizing and converting network-on-chip (NoC) devices, wherein the tool comprises: Conversion module and The system comprises an optimization module that communicates with the aforementioned conversion module, The aforementioned tool is The system receives information about the NoC, including multiple links, multiple switches, and multiple constraints, as input. By using the conversion module to cluster at least two links selected from the plurality of links, a link cluster that conforms to at least one of the plurality of constraints is generated. Using the optimization module, a new link is generated to replace the link cluster by collapsing the link cluster. By using the conversion module to cluster at least two switches selected from the plurality of switches, a first potential switch cluster is formed that satisfies at least one of the plurality of constraints. Using the optimization module, a folding switch is generated by folding the first potential switch cluster. A tool that generates a cycle-free conversion No.C based on the No.C, using the aforementioned new link and the folding switch.
2. The aforementioned tool is Using the conversion module, determine whether other links can be added to the link cluster by traversing the plurality of links adjacent to the link cluster. If it is determined that the other links can be added to the link cluster, the other links are added to the link cluster. The tool according to claim 1, wherein a second link cluster for replacing the link cluster is generated by collapsing the link cluster and the other links using the optimization module.
3. The aforementioned tool is Using the aforementioned conversion module, it is determined whether or not other switches can be added to the folding switch by traversing the plurality of switches adjacent to the folding switch. If it is determined that the aforementioned other switch can be added to the folding switch, a second folding switch is generated by adding the aforementioned other switch to the folding switch. The tool according to claim 1, wherein the optimization module is used to generate a second conversion No. C based on the conversion No. C using the second folding switch.
4. The tool collapses the first potential switch cluster. Selecting at least two of the aforementioned switches, Removing at least two of the aforementioned switches, This includes adding new switches to replace the at least two switches that have been removed, The tool according to claim 1, wherein the new switch provides connectivity for the at least two switches.
5. The aforementioned No.C is a cycle-free network, The tool according to claim 1, wherein the conversion No. C is a cycle-free network.
6. The tool clusters the at least two links, Creating an empty link cluster, Selecting the first link from the aforementioned multiple links, This includes selecting a second link from the plurality of links, wherein the first link and the second link form a link set. The tool according to claim 1, comprising generating the link cluster by assigning the link set to the empty link cluster.
7. The second link is selected if it has features in common with the first link. The tool according to claim 6, wherein the aforementioned features include the same direction and proximity.
8. The tool according to claim 1, wherein the tool sorts a list of link clusters in descending order of gain using a gain function, and the link cluster with the highest gain from the list of link clusters is processed first.
9. The tool according to claim 1, wherein the tool determines whether collapsing the link cluster introduces a topology loop using network cycling.
10. The tool according to claim 1, wherein the tool uses network cycle exploration to determine whether a potential switch cluster introduces a topology loop.