System and method for parallel algorithm and scalable architecture for routing in beneš networks
Patent Information
- Authority / Receiving Office
- IL · IL
- Patent Type
- Applications
- Current Assignee / Owner
- RAMOT AT TEL AVIV UNIVERSITY LTD
- Filing Date
- 2024-12-19
- Publication Date
- 2026-08-01
AI Technical Summary
Current optical circuit switches (OCSs) face challenges such as long switching times in micro electromechanical systems (MEMS) and high costs associated with arrayed waveguide grating routers (AWGRs), which hinder their adoption in data switches.
A system and method for routing in Benes networks using a parallel algorithm and scalable architecture, which supports full and partial input permutations. The algorithm limits processing time to ((2^N)^2) steps by potentially forfeiting a few input demands, achieving close to 100% utilization and reducing connectivity complexity to (2^N), a significant improvement over previous solutions.
The proposed solution achieves close to 100% utilization for both full and partial input permutations, reduces connectivity complexity, and scales more efficiently, addressing the limitations of existing OCS technologies.
Smart Images

Figure 00000044_0000 
Figure 00000045_0000 
Figure 00000045_0001
Abstract
Description
Attorney Ref. FRAMO-P036-WO System and Method for Parallel Algorithm and Scalable Architecture for Routing in Beneš Networks CROSS-REFERENCE TO RELATED APPLICATION
[0001] This application is related to and claims priority under 35 U.S.C. § 119(e) to U.S. Patent Application No.63 / 614,458 filed December 22, 2023, entitled “Parallel Algorithm and Scalable Architecture for Routing in Beneš Networks,” the entire contents of which is incorporated herein by reference. BACKGROUND
[0002] The gap between the rapid growth of data center traffic and slower growth in Electrical Packet Switches (EPSs) capacity is increasing and is expected to be even worse as EPSs nearly reach their limit in terms of capacity and power consumption, while cloud traffic roughly doubles every year.
[0003] EPSs require optical-electrical conversion and increased number of SERDES lanes with the increase of switch capacity. High-radix EPSs have scalability issues due to high power consumption, large Silicon size, and cost. Furthermore, as Moore’s law approaches its limits, moving beyond 5nm CMOS technology will be extremely challenging. To resolve the electrical bottleneck of power, size, and cost, optical circuit switches (OCSs) are considered as a replacement for EPSs in the network. OCSs are transparent to data rates, they can have similar or higher radix than typical EPSs, they are typically used as a crossbar without any packet processing or buffering and hence consume much less power than EPSs, and with OCSs there is no need to convert back and forth from electrical signals to light.
[0004] However, many of the popular technologies used to build OCS suffer from significant drawbacks, which hinders their usage for data switches. In particular, micro electromechanical systems (MEMS) have long switching times, while arrayed waveguide grating routers (AWGRs) require expensive multi-color lasers.
[0005] It is with these issues in mind, among others, that various aspects of the disclosure were conceived. SUMMARY
[0006] The present disclosure is directed to a system and method for routing in Benes networks using a parallel algorithm and scalable architecture. In addition, the present disclosure is directed to a new routing algorithm for Beneš networks combined with a scalable hardwareAttorney Ref. FRAMO-P036-WO architecture that supports full and partial input permutations. The processing time of the algorithm is limited to ^^((^^^^^^2^^)2steps (iterations) by potentially forfeiting routing of a few input demands; however achieves close to 100% utilization for both full and partial input permutations. The algorithm and architecture allow a reduction of the connectivity complexity to ^^(^^2), a log ^^ improvement over previous solutions.
[0007] In one example, a method for routing a Beneš network may include for a plurality of processing elements that represent a plurality of switching elements in a Beneš network and a communication bus in between the plurality of processing elements, wherein two sub-networks are disposed in between a first column of the Beneš network and a last column of the Beneš network, wherein the first column comprises a first set of switching elements in the plurality of switching elements, and wherein the last column comprises a second set of switching elements in the plurality of switching elements setting an initial state for the plurality of processing elements that represent switching elements of the first column, configuring the plurality of processing elements that represent the switching elements of the first column to communicate in a communication with each other to detect a plurality of collisions and to resolve the plurality of collisions by flipping states in a plurality of iterations, the communication performed by a centralized collision bus, and when the plurality of iterations are completed, resolving any remaining collisions, if any, by forfeiting an input demand from one colliding input port per a remaining collision, resulting in a reduction of utilization, wherein each bit in the centralized collision bus represents a switching element in the second set of switching elements.
[0008] In another example, a method for routing a Beneš network may include distributing a plurality of demands to a plurality of processing elements in the Beneš network, setting an initial state for each of the plurality of processing elements, performing an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network, forfeiting, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision, setting states of a plurality of switching elements in a last column of the Beneš network, delivering at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network, and either (i) informing, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) informing an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the method is performedAttorney Ref. FRAMO-P036-WO recursively and repeatedly for each of the one or more sub-networks in parallel, ending in the smallest sub-network of the one or more sub-networks.
[0009] In another example, a system for routing a Beneš network may include a plurality of processing elements that represent a plurality of switching elements in a Beneš network and a communication bus in between the plurality of processing elements, wherein two sub-networks are disposed in between a first column of the Beneš network and a last column of the Beneš network, wherein the first column comprises a first set of switching elements in the plurality of switching elements, and wherein the last column comprises a second set of switching elements in the plurality of switching elements, the plurality of processing elements setting an initial state for the plurality of processing elements that represent switching elements of the first column, configuring the plurality of processing elements that represent the switching elements of the first column to communicate in a communication with each other to detect a plurality of collisions and to resolve the plurality of collisions by flipping states in a plurality of iterations, the communication performed by a centralized collision bus, and when the plurality of iterations are completed, resolving any remaining collisions, if any, by forfeiting an input demand from one colliding input port per a remaining collision, resulting in a reduction of utilization, wherein each bit in the centralized collision bus represents a switching element in the second set of switching elements.
[0010] In another example, system for routing a Beneš network may include a plurality of processing elements to receive a plurality of demands in the Beneš network, set an initial state for each of the plurality of processing elements, perform an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network, forfeit, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision, set states of a plurality of switching elements in a last column of the Beneš network, deliver at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network, and either (i) inform, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) inform an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the system is a recursive system to repeatedly for each of the one or more sub- networks in parallel, end in the smallest sub-network of the one or more sub-networks using recursion.Attorney Ref. FRAMO-P036-WO
[0011] In another example, a non-transitory computer-readable storage medium may have instructions stored thereon that, when executed by at least one computing device cause the at least one computing device to perform operations including distributing a plurality of demands to a plurality of processing elements in the Beneš network, setting an initial state for each of the plurality of processing elements, performing an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network, forfeiting, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision, setting states of a plurality of switching elements in a last column of the Beneš network, delivering at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network, and either (i) informing, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) informing an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the method is performed recursively and repeatedly for each of the one or more sub-networks in parallel, ending in the smallest sub-network of the one or more sub-networks.
[0012] These and other aspects, features, and benefits of the present disclosure will become apparent from the following detailed written description of the preferred embodiments and aspects taken in conjunction with the following drawings, although variations and modifications thereto may be effected without departing from the spirit and scope of the novel concepts of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings illustrate embodiments and / or aspects of the disclosure and, together with the written description, serve to explain the principles of the disclosure. Wherever possible, the same reference numbers are used throughout the drawings to refer to the same or like elements of an embodiment, and wherein:
[0014] Figure 1 shows a recursive construction of an NxN Beneš network according to an example of the instant disclosure.
[0015] Figure 2A and Figure 2B show average utilization for basic and improved method and a K-collisions probability distribution according to an example of the instant disclosure.
[0016] Figure 3 shows a collision bus implemented with wired-OR according to an example of the instant disclosure.
[0017] Figure 4 illustrates collision bus steps for N / 2 bits and N bits according to an example of the instant disclosure.Attorney Ref. FRAMO-P036-WO
[0018] Figure 5 shows an example of step two of a Beneš network routing algorithm according to an example of the instant disclosure.
[0019] Figure 6 shows a single iteration pipeline according to an example of the instant disclosure.
[0020] Figure 7 illustrates first column average utilization according to an example of the instant disclosure.
[0021] Figure 8 shows histograms of a number of iterations with 100% load and 90% load according to an example of the instant disclosure.
[0022] Figure 9 shows average utilization of an NxN network, fixed mode for 100% input load and 90% input load according to an example of the instant disclosure.
[0023] Figure 10 shows distribution of utilization for 100% input demands for twelve iterations per layer according to an example of the instant disclosure.
[0024] Figure 11 shows a method of routing a Beneš network according to an example of the instant disclosure.
[0025] Figure 12 shows another method of routing a Beneš network according to an example of the instant disclosure.
[0026] Figure 13 shows an example of a system for implementing certain aspects of the present technology. DETAILED DESCRIPTION
[0027] The present disclosure is more fully described below with reference to the accompanying figures. The following description is exemplary in that several embodiments are described (e.g., by use of the terms “preferably,” “for example,” or “in one embodiment”); however, such should not be viewed as limiting or as setting forth the only embodiments of the present disclosure, as the disclosure encompasses other embodiments not specifically recited in this description, including alternatives, modifications, and equivalents within the spirit and scope of the invention. Further, the use of the terms “invention,” “present invention,” “embodiment,” and similar terms throughout the description are used broadly and not intended to mean that the invention requires, or is limited to, any particular aspect being described or that such description is the only manner in which the invention may be made or used. Additionally, the invention may be described in the context of specific applications; however, the invention may be used in a variety of applications not specifically described.Attorney Ref. FRAMO-P036-WO
[0028] The embodiment(s) described, and references in the specification to “one embodiment”, “an embodiment”, “an example embodiment”, etc., indicate that the embodiment(s) described may include a particular feature, structure, or characteristic. Such phrases are not necessarily referring to the same embodiment. When a particular feature, structure, or characteristic is described in connection with an embodiment, persons skilled in the art may effect such feature, structure, or characteristic in connection with other embodiments whether or not explicitly described.
[0029] In the several figures, like reference numerals may be used for like elements having like functions even in different drawings. The embodiments described, and their detailed construction and elements, are merely provided to assist in a comprehensive understanding of the invention. Thus, it is apparent that the present invention can be carried out in a variety of ways, and does not require any of the specific features described herein. Also, well-known functions or constructions are not described in detail since they would obscure the invention with unnecessary detail. Any signal arrows in the drawings / figures should be considered only as exemplary, and not limiting, unless otherwise specifically noted. Further, the description is not to be taken in a limiting sense, but is made merely for the purpose of illustrating the general principles of the invention, since the scope of the invention is best defined by the appended claims.
[0030] It will be understood that, although the terms first, second, etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. Purely as a non-limiting example, a first element could be termed a second element, and, similarly, a second element could be termed a first element, without departing from the scope of example embodiments. As used herein, the term "and / or" includes any and all combinations of one or more of the associated listed items. As used herein, the singular forms "a", "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be noted that, in some alternative implementations, the functions and / or acts noted may occur out of the order as represented in at least one of the several figures. Purely as a non-limiting example, two figures shown in succession may in fact be executed substantially concurrently or may sometimes be executed in the reverse order, depending upon the functionality and / or acts described or depicted.
[0031] Conditional language, such as, among others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, isAttorney Ref. FRAMO-P036-WO generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements and / or steps. Thus, such conditional language is not generally intended to imply that features, elements and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements and / or steps are included or are to be performed in any particular embodiment.
[0032] Beneš / CLOS architectures are common scalable interconnection networks widely used in backbone routers, data centers, on-chip networks, multi-processor systems, and parallel computers. Recent advances in Silicon Photonic technology, especially MZI technology, have made Beneš networks a very attractive scalable architecture for optical circuit switches.
[0033] Numerous routing algorithms for Beneš networks were developed starting with linear algorithms having time complexity of ^^(^^^^^^^^2^^) steps. Parallel routing algorithms were developed to satisfy the stringent timing requirements of high-performance switching networks and have time complexity of ^^((^^^^^^2^^)2. However, their implementation requires ^^(^^2^^^^^^2^^) wires (termed connectivity complexity), and thus are difficult to scale.
[0034] This disclosure discusses a new routing algorithm for Beneš networks combined with a scalable hardware architecture that supports full and partial input permutations. The processing time of the algorithm is limited to ^^((^^^^^^2^^)2steps (iterations) by potentially forfeiting routing of a few input demands; however achieves close to 100% utilization for both full and partial input permutations. The algorithm and architecture allow a reduction of the connectivity complexity to ^^(^^2), a log ^^ improvement over previous solutions.
[0035] This disclosure proves the algorithm correctness, and analyzes its performance analytically and with large scale simulations.
[0036] INTRODUCTION
[0037] Crossbar networks, the simplest solution for parallel switches, contain ^^2crosspoints (for ^^ × ^^ switch) and thus are too costly for large ^^. A three-layer network ofcrossbars can reduce the cost of the switch to only 2√2^^3 / 2.
[0038] Indeed, CLOS topologies have become a standard in large scale routers.
[0039] Beneš networks are a recursive construction of a CLOS network, which result in a multi-stage switching network. It uses 2x2 switching elements (SEs) arranged in a matrix of^^ / 2 columns by 2 ^^^^^^2(N) − 1 rows producing a rearrangeable nonblocking multistageAttorney Ref. FRAMO-P036-WO interconnection network. ’Routing’ in a Beneš network refers to the process of producing the state setting of all the SEs in the Beneš network in order to satisfy an input–output routing demand.
[0040] The gap between the rapid growth of data center traffic and slower growth in Electrical Packet Switches (EPSs) capacity is increasing and is expected to be even worse as EPSs nearly reach their limit in terms of capacity and power consumption, while cloud traffic roughly doubles every year.
[0041] EPSs require optical-electrical conversion and increased number of SERDES lanes with the increase of switch capacity. High-radix EPSs have scalability issues due to high power consumption, large silicon size, and cost. Furthermore, as Moore’s law approaches its limits, moving beyond 5nm CMOS technology will be extremely challenging. To resolve the electrical bottleneck of power, size, and cost, optical circuit switches (OCSs) are considered as a replacement for EPSs in the network. OCSs are transparent to data rates, they can have similar or higher radix than typical EPSs, they are typically used as a crossbar without any packet processing or buffering and hence consume much less power than EPSs, and with OCSs there is no need to convert back and forth from electrical signals to light.
[0042] However, many of the popular technologies used to build OCS suffer from significant drawbacks, which hinders their usage for data switches. In particular, micro electromechanical systems (MEMS) have long switching times, while arrayed waveguide grating routers (AWGRs) require expensive multi-color lasers.
[0043] Mach–Zehnder interferometer (MZI) is an attractive technology for OCS. Its basic building block is a 2x2 switch that is naturally used to build Beneš networks. Beneš networks are also an attractive architecture for networks-on-chip. In 2014, an experimental recirculating loop through a 2x2 switch, equivalent to a propagation through a 128x128 Beneš network, has been demonstrated. In 2016, a 16x16, 32x32 and 64x64 MZI based Beneš network OCSs were demonstrated. Recently in 2021, capability of building a 256x256 Beneš network based OCS was analyzed and evaluated.
[0044] The fast switching time (a few nano-seconds) of MZI switches facilitates large high-speed switches, but Beneš networks suffer from routing algorithms that are difficult to scale largely due to the significant amount of wiring required to connect the routing processing elements that are used by today’s algorithms. The connectivity complexity, which is the number of wires required to connect the switching element control, had only recently emerged as an importantAttorney Ref. FRAMO-P036-WO obstacle in scaling Beneš networks, since the wires occupy significant space. This disclosure suggests a significant improvement in the connectivity complexity, and thus removes a significant barrier in the scaling of Beneš networks.
[0045] The first Beneš routing algorithms were serial and as a result required at least ^^(^^^^^^^^2^^) operations to complete. This makes them impractical for large switches.
[0046] In the early 1980s, two teams suggested simultaneously parallel routing algorithm with a time complexity of ^^((^^^^^^2^^)2). Both algorithms used Completely Interconnected Computers (CIC) topology, and were not capable of handling partial permutations. Since the CIC topology requires all PEs (Processing Elements) to communicate with each other, with at least the PE ID, requiring log2 N bits, the connectivity complexity is ^^(^^2^^^^^^2^^)for each layer.
[0047] In 1995, Lee and Oruc proposed a solution for partial permutations by matching idle input port to idle output port, which requires significant pre-processing and complicates the data structures. In 1996, Lee and Liew published a parallel algorithm running in ^^((^^^^^^2^^)2) time complexity using ^^ / 2 processing elements that can also handle partial permutations. In 2009, Kai et al. developed a design based on Lee and Liew’s algorithm that reduces the hardware complexity from ^^(^^2) to ^^(^^(^^^^^^2^^)2) and time complexity from ^^((^^^^^^2^^)2) to^^(log ^^)using pipelined architecture, but their solution cannot handle partial permutations.
[0048] In 2016, Jiang et al. implemented Lee and Liew’s algorithm using a centralized memory for small ^^ values, ^^ ∈ [8, 32], with time complexity of ^^((^^^^^^^^)2) and connectivity complexity of ^^(^^2^^^^^^2^^)having a centralized architecture where each PE has a dedicated interface to a centralized control logic and a centralized memory.
[0049] Finally, in 2021, Koloko at al. developed yet another parallel implementation to Lee and Liew’s algorithm having time complexity of ^^((^^^^^^2^^)2) and hardware complexity of ^^(^^(^^^^^^2^^)2) since each PE communicates with every other PE the connectivity complexity remains ^^(^^2^^^^^^2^^).
[0050] This disclosure discusses an efficient, functional, fast and, most importantly, scalable parallel algorithm to route a Beneš network. The algorithm works iteratively, where each iteration potentially increases the number of routed input demands to their corresponding outputs hence allowing a trade-off between processing time and number of connected inputs to outputs. The algorithm performs better given partial permutations.Attorney Ref. FRAMO-P036-WO
[0051] As a trade-off can be made between processing time and the number of serviced input demands of a given input permutation, it is possible to stop processing at a time complexity of ^^((^^^^^^2^^)2) and achieve, on average, close to 100% success to route input demand (more specifically ≥ 95% for small size network and ≥92% for higher network size). As discussed herein, connectivity complexity is only ^^(^^2)wires an improvement of ^^(^^^^^^2^^) over previous solutions, which makes the algorithm easier to scale to high-radix switches.
[0052] One idea that allows the system described herein to break away from the wiring scalability challenge is a design of a ’collision’ bus. Using this bus, the disclosure presents an efficient algorithm to route almost 100% of a demand in just ^^((^^^^^^2^^)2) steps.
[0053] The disclosure further provides the theoretical background as the foundations for the algorithm. The disclosure further describes and defines the interfaces between PEs. The disclosure then presents the routing algorithm. Additionally, the disclosure presents simulation results.
[0054] PRELIMINARIES
[0055] In this section the disclosure presents the theoretical background of the routing algorithm, the utilization of the first column of Beneš network for random setting of first column SEs’ states and finally, the routing principles.
[0056] The proofs for all lemmas are given herein.
[0057] Definitions, Propositions and Lemmas
[0058] The disclosure presents properties of the Beneš network that are the basis of this disclosure’s routing algorithm. All letter variables signify integer values for SEs, PEs and ports.
[0059] As shown in Figure 1, a minimal re-arrangeable non-blocking NxN Beneš network, is recursively built as follows:
[0060] 1) A first column of ^^ / 2 2 × 2 switching elements (SEs)
[0061] 2) A last column of ^^ / 2 2 × 2 SEs
[0062] 3) Middle stage comprised of two ^^ / 2 × ^^ / 2 sub-networks,
[0063] upper and lower, that are built recursively
[0064] 4) Recursion ends at a single column of 2 × 2 SEs
[0065] Figure 1 shows a Recursive construction of NxN Beneš network 100 according to an example of the instant disclosure.Attorney Ref. FRAMO-P036-WO
[0066] Every SE in the first column has one output connected to the upper sub-network in the middle stage and one output to the lower sub-network. The inputs to the SEs of the last column are symmetrically connected to the middle stage.
[0067] The entire network can be viewed as a matrix of SEs having 2 ^^^^^^2 N − 1 columnsand N / 2 rows where there are ^^ connections between each column. SEs in each column areenumerated from 0 (top) to N / 2 − 1 (bottom), and the ports of ^^^^^^ are denoted by 2^^ and 2^^ +1. Definition 1. A SE can be in either of two states:
[0069] 0) A ’bar’ state (= or 0) where ^^0 ← ^^0 and ^^1 ← ^^1
[0070] 1) A ’cross’ state (X or 1) where ^^0 ← ^^1 and ^^1 ← ^^0
[0071] Definition 2. If ^^^^^^is in bar state, port 2k is connected to the upper sub-network and port 2k + 1 to the lower sub-network and the opposite for cross state. The definition is symmetrically applied to the last column SEs.
[0072] Given a ^^ × ^^ network, let ^^ be the set of inputs and ^^ be the set outputs, where^^ = ^^ = {0, 1, 2, ... , ^^ − 1}. A function π : I→O is called a permutation if it is bijective(namely, it is one-to-one and onto). It is possible to widen the definition of a permutation byincluding the case where an input is idle, which is denoted by π(^^) = χ for an idle input ^^.
[0073] Definition 3. A function π : I→O ∪ {x} is called a permutation if π(i) = π(j) ∈ O ⇒ i = j.
[0074] The inputs to ^^^^^^are ^^^^0= π(2k) and ^^^^1= π(2k + 1). Using Definition 3, if ^^^^^^state is bar, then ^^^^0= π(2k) and ^^^^1= π(2k + 1) and the opposite if ^^^^^^state is cross.
[0075] An example of an input permutation is given in Eq. (1) for a 8x8 Beneš network. The upper row in the matrix is the ordered input ports from 0 to N-1 and the lower row is the outputs requested by the corresponding inputs. In this example, input 0 requests route to output 2, input 1 to 3, input 2 is idle, and so on, namely, π(0) = 2, π(1) = 4, π(2) = x.
[0076] ^^ = (01234567 ) (1) ^^ 53017
[0077] text the disclosure uses the double division symbol ( / / ) to mark an ∆ integer division, namely taking the integer portion of the result: i / / j =→ ⌊^^ / ^^ ⌋ . Also π(i) is calledan input demand, or simply demand, for input port i. The lemmas below are unique to the algorithm.Attorney Ref. FRAMO-P036-WO
[0078] Definition 4. A collision between two input ports, i and j, of different SEs occurs when they both have demands to the same SE in the last column, that is π(i) / / 2 = π(j) / / 2, and both are routed through the same sub-network. The SE in the last column has only one input from each sub-network therefore it cannot accept two demands from the same sub-network.
[0079] For example, when applying the permutation shown in Eq. (1) to a 8 × 8 Benešnetwork and setting all SEs in the first column to bar state, following collisions are created:
[0080] Input ports 0 and 4 collide through upper sub-network
[0081] Input ports 1 and 3 collide through lower sub-network
[0082] A different state setting of the SEs in the first column, such as{^^^^0, ^^^^1, ^^^^2, ^^^^3} = {1, 0, 0, 0} , resolves both collisions.
[0083] Lemma 1. A collision occurs between exactly two demands
[0084] Proposition 1. A single SE in the first column can be involved in at most two collisions, one in the upper sub-network and the other in the lower sub-network.
[0085] Proposition 2. A collision occurs only between odd input demand and even input demand where the corresponding input ports are connected to different SEs in the first column, collision cannot occur between two even input demands or two odd input demands.
[0086] Lemma 2. Given 2 colliding demands (^^^^^^and ^^^^^^in first column, where ^^ ≠^^ , i is an input port^^ ^^ ^^demands are routed through the same sub-network), then flipping the state of either one SE (but not both) resolves this collision but may create a new collision in the same sub-network.
[0087] Lemma 3. Flipping a state of a first column SE that has two collisions, one in each sub-network, resolves both collisions, but may create a new collision in either of the sub-networks. This means the total number of collisions cannot be increased.
[0088] Corollary 1. Flipping the state of a single SE in the first column having colliding input demand(s) cannot increase the total number of collisions.
[0089] Following Lemmas 2 and 3, the process of flipping the state of one colliding SE, in each collision, cannot increase the total number of collisions, and thus produces a monotonically nonincreasing number of collisions after each flip.
[0090] Proposition 3. Given an input port i with a demand to output port j, that is j = π(i), the state of ^^^^^^(^^) / / 2in the last column is set as follows:
[0091] Cross if π(i) is odd and routed to upper sub-networkAttorney Ref. FRAMO-P036-WO
[0092] Cross if π(i) is even and routed to lower sub-network
[0093] Bar for all other cases
[0094] Lemma 4. If there are no collisions between first column SEs, collisions between last column SEs are not possible.
[0095] B. First column SEs initial state and collision probability
[0096] In this section the disclosure discusses the probability for collisions given a random initial setting of states of SEs in the first column and calculate the utilization of Beneš network. Then the disclosure presents initial setting heuristic that improves the utilization, where the term ’Utilization’ is defined as follows:
[0097] Definition 5. Utilization is equal to the number of established routes from inputs to outputs, based on the input demands that the algorithm achieved, divided by the number of valid input demands. A utilization of 100% means that routes were established for all input demands.
[0098] Probability of collisions:
[0099] Observation 1. In a full input permutation (no input port is idle) there are exactly N / 2 input ports that are connected to the upper (and to lower) sub-network for any given SEs’ states.
[0100] Lemma 5. Based on observation 1, in a full input permutation, a collision in one sub-network always implies a collision in the other sub-network as well.
[0101] The number of collisions, if any, for a full input permutation demand is always even (Lemma 5). Therefore, it is sufficient to only count the number of collisions in the upper sub- network due to this symmetry. When the input permutation is not full, the total number of collisions can be odd.
[0102] Corollary 2. The maximum number of collisions for NxN Beneš network is N / 4 for each sub-network, that is a total of N / 2 collision for the entire network.
[0103] The disclosure determines the probability of a collision between SEs in first column to one sub-network given full input permutation where the total number of collisions is exactly twice.
[0104] It is possible to first calculate the total number of permutations for N / 2 ports in one sub-network. The first input port selects 1 out of N output ports to be routed to, second input port remains with 1 out of N − 1 output ports, ^^^^ℎoutput port remain with 1 out of N − (k − 1) ports, and so on. The number of ways to connect N / 2 input ports to one sub-network is:Attorney Ref. FRAMO-P036-WO ^^
[0105] ^^ ∙ (^^ − 1) … (^^ −^^ 2+ 1) = Π2−1 ^^− 0(^^ − ^^) =^^! ^^ ( 2 (2) 2)! the. second pair has N / 2 − 2 options and so on. As the ordering between the selected pairs is irrelevant, divide by ^^!, Thus: 1 ^^ ^^ ^^ ^^ ^^!∙ (( 2 2 −2 2 −2 (^^−1) 1 2)!
[0107] 2) ∙ (2) … (2) =^^! ∙ ^ 2^^^ ∙( ! (3) Eachan leaving two less ports to select from. A colliding pair of input ports are routed to the two output ports and a single non-colliding input port must not allow the other output port of the same SE to be selected by others. In total there are N / 2− k selections where the first has N ports to choose from, the second has N / 2 ports to choose from and so on. This count is depicted in Eq. (4). ^^ − ^^
[0109] ^^ ∙ (^^ − 2 ) … (^^ − 2 (^^− 1 − ^^)) =22 ^^∙( 2)! (4)Prob(N, k) is determined, the probability of k-collisions in a single sub-network of a ^^ × ^^ Benešnetwork when setting SEs in the first column to random states (bar or cross) for a full input permutation demand, Eq. (5). ^^ 3) ^^ ^^^^ ^^ =[( −2^^ ^^^^2)!] ∙ 22 ^^ (5)having no collisions. This is done by allowing the first input port to choose 1 out of N output ports, the second chooses 1 out of N − 2 ports (to avoid collision), the third chooses 1 out of N − 4 ports and so on for N / 2 input ports. ^^
[0113] ^^ ∙ (^^ − 2 ) ∙ (^^ − 4) … 2 = 2 ∙ (^^ 2 2) ! (6)
[0114] The probability for 0 collisions is obtained by dividing expression (6) by (2), which is the same as setting k = 0 in Eq. (5). When simplifying the result, by applying Stirling’s approximation to N! and (N / 2)!, Eq. (7) is obtained.Attorney Ref. FRAMO-P036-WO ^^ ^^
[0115] ^^^^^^^^ (^^, 0) =22∙ (22)! ^^^^ ^^!≈ √2^^−1(7)input ports, resulting in reduced utilization. The average utilization of the first column when the SEs’ states are set arbitrarily given 100% input load is derived by multiplying Eq. (5) by the remaining input demands after forfeiting colliding demands and summarize for all collisions, shown in Eq. (8). ^^
[0118] ^^1 (^^) =1 ^^∑4 ^^=0^^^^^^^^ (^^, ^^) ∙ (^^ − 2^^)(8)converges to distributions of k-collisions for N=32, 64, 128, 256 network sizes.
[0120] The maximum probability of collision occurs for k = N / 8. Each sub-network has N / 2 ports hence at most N / 4 collisions.
[0121] Figure 2A shows average utilization for basic and improved methods.
[0122] Figure 2B shows K-collisions Probability distribution.
[0123] Due to symmetry the maximum probability of collision relies at half of the total number of collisions, that is N / 8. Another observation is that the maximum probability of collision (k = N / 8) as a function of N decreases in a logarithmic scale.
[0124] The lower bound of the average utilization for the entire ^^ × ^^ Beneš network,when SEs in all layers are set to a random state, is the multiplication of ^^1for all sub-networks, Eq. (9).
[0125] ^^^^^^^^^^^^ ^^^^^^^^^^ (^^) = ∐log2^^ ^^=2 ^^1(2^^) (9)are reduced resulting in less probability for collisions; therefore the average utilization increasesfor the sub-networks. For example, the calculated lower bound utilization for 32 × 32 Benešnetwork is 38.05% however, simulations show an average utilization of 65.5%.
[0127] Improved initial state of the first column:Attorney Ref. FRAMO-P036-WO
[0128] No collisions occur if all demands to even output ports are directed to the upper sub-network and all demands to odd output ports are directed to the lower sub-network (or the opposite). Based on that, the disclosure introduced a heuristic:
[0129] For each ^^^^^^, if input ^^0is odd and input ^^1is even, set SE state to cross, otherwise set SE state to bar.
[0130] Simulations show that this heuristic increases the average utilization of the first column by additional ∼6% for N > 32, shown in Figure 2A in the line marked improved 204.
[0131] C. Routing Principles
[0132] Parallel Beneš network routing algorithms calculate states of SEs in the first and last columns and set demands to the sub-networks. This process is repeated recursively in each subnetwork until all layers are routed.
[0133] The routing algorithm as discussed herein has the ability to tradeoff between completion time of calculating the states of the first columns and utilization (which can be below 100%) which allows for easy scalability.
[0134] The routing algorithm as discussed herein is based on Corollary 1 and the fact that it is enough for a PE to know that it has a colliding demand without knowing which is the other colliding PE.
[0135] It is possible to refer to ^^^^^^as the processing element of ^^^^^^in first and in last columns.
[0136] To resolve the states of the first column SEs, the algorithm operates in iterations. In each iteration, the PEs communicate with each other to detect and resolve collisions by flipping SEs’ states. Once the iteration process is completed, the states of the last column SEs and demands to the sub-networks are calculated. At the end of the recursion, the innermost subnetworks indicate to the outmost layer about forfeited input demands due to collisions.
[0137] INTERFACES BETWEEN PES
[0138] PEs communicate between them in order to detect and resolve collisions.
[0139] Collision bus
[0140] A centralized bus, named ”Collision bus”, is used for the communication between PEs. Each bit in the bus represents a required SE in the last column, therefore the size of the bus is N / 2 and the collision bus is an input to all PE.Attorney Ref. FRAMO-P036-WO
[0141] Each PE has N / 2 bits output bus. ^^^^^^can set bus bits π(2k) / / 2 or π(2k+1) / / 2 indicating the index of the required SE in the last column, based on its valid input demands, π(2k), π(2k + 1). An idle input demand cannot set any bit.
[0142] Collision bus bit [i] is the result of an OR gate between bit [i] of output busses of all N / 2 PEs. Wired-OR technology can eliminate the use of actual OR gates. The implementation of the collision bus based on wired-OR 300 is shown in Figure 3.
[0143] In total, each PE has N interface bits, N / 2 output bits and N / 2 input bits, and the total wires for N / 2 PEs is ^^2 / 2 , that is O(N2), a reduction of log2^^ of existing solutions.
[0144] Figure 3 shows a collision bus implemented with wired-OR 300 according to an example of the instant disclosure.
[0145] Proposition 4. Based on Lemma 1 and proposition 2, exactly 2 ports, i and j, having demands π(i) and π(j) from 2 different PEs, ^^^^^^and ^^^^^^, can collide. Therefore, if bit [π(i) / / 2] on the collision bus is set by ^^^^^^, only ^^^^^^relates to that bit and, π(i) and π(j) are not both odd or both even.
[0146] Proposition 5. Based on Proposition 4, when all PEs with demands to even output ports routed through the upper subnetwork set their corresponding bits on the collision bus, each bit on the bus can be set by one and only one such PE.
[0147] A set bit on the bus can match a demand of one and only one PE that has demand to an odd output port that is routed through the upper sub-network
[0148] When ^^^^^^sets bit [d] on its output bus, it means that ^^^^^^’publishes’ demand to ^^^^^^in the last column. Based on proposition 5, when all PEs with even (or odd) output demands to upper (or lower) sub-network publish their demands, a single bit on the collision bus can be set by only one PE. It is not possible for two such PEs to set the same bit.
[0149] This property can only be achieved if PEs are divided to sets of publishing and listening PEs for the same sub-network, set 1 publish and set 2 listen or the opposite and set 3 publish and set 4 listen or the opposite. Sets are as follows:
[0150] 1) PEs with even demand routed via upper sub-network
[0151] 2) PEs with odd demand routed via upper sub-network
[0152] 3) PEs with even demand routed via lower sub-network
[0153] 4) PEs with odd demand routed via lower sub-network
[0154] An IterationAttorney Ref. FRAMO-P036-WO
[0155] Following are iteration steps to detect / resolve collisions:
[0156] Publish demands: PEs of one set publish their required SE in last column on the collision bus (that is demand / / 2).
[0157] Listen and detect collision: Each PE in an opposite set listens to a bit in the collision bus, corresponding to its required SE in the last column. If the bit is set, a collision is detected by the listener PE.
[0158] Flip state: Listening PEs having two collisions flips their state to resolve the collision, if having one collision, they flip their state with probability 1 / 2.
[0159] Flip message: PEs that decided not to flip their state send a similar message to ’publish’, on the collision bus. The publishing PEs listen to the bus and flip their state if their demand is equal to the set bit on the bus.
[0160] A single iteration, shown in Figure 4, contains steps 1,2 for the upper sub-network (sets 1 and 2), then repeat steps 1,2 for the lower sub-network (sets 3 and 4), then steps 3,4.
[0161] An iteration can be reduced to 2 messaging phases by doubling the collision bus width to N bits, shown in Figure 4, such that two publish phases are performed in one message where each uses half of the bus. Selecting N / 2 or N bits bus depends on implementation constraints.
[0162] Figure 4 shows collision bus steps 400 for (a) N / 2 bits, (b) N bits according to an example of the instant disclosure.
[0163] Forfeit bus
[0164] Input demands may be forfeited in any sub-network due to unresolved collisions following the iterations process. The forfeited input demands must be propagated to the outmost layer. For that purpose, input demands are propagated in full to all sub-layers, during the recursive iterative process, even if they are forfeited in some middle layer. At the end of the recursion, the sub-networks in the innermost layer contains all valid, invalid or forfeited input demands, however without the knowledge of their corresponding input port.
[0165] To pass the information of invalid (forfeited) demands from innermost to outmost layer, a dedicated bus, named “Forfeit bus,” is used from innermost layer PEs to outmost layer PEs. The structure of this bus is the same as the collision bus, it contains N bits where bit [i] corresponds to input demand [i].
[0166] Each PE in the innermost layer has N outputs to the Forfeit bus corresponding to N input demands. Each PE can set up to two bits on this bus corresponding to its invalid / forfeitedAttorney Ref. FRAMO-P036-WO input demands. All N bits from all innermost PEs are ORed together, similar to the collision bus, to create the forfeit bus.
[0167] Each PE in the outmost layer listens to the forfeit bus and checks if its valid input demands are set on the bus, if so, the PE invalidate these input demands.
[0168] ALGORITHM
[0169] A Beneš network routing algorithm contributes: (1) a new iterative approach to resolve routing conflicts; (2) reduces connectivity between PEs; and (3) fast average completion time of ^^ ((log2^^)2).
[0170] This is the first Beneš routing algorithm that trades-off processing time and utilization.
[0171] The routing algorithm contains the following six steps:
[0172] 1) Distribute demand and Set Initial state
[0173] 2) Modify states of SEs in first column by performing iterative collision detection and resolution process
[0174] 3) Remove demands of remaining colliding inputs
[0175] 4) Set switching states of SEs in last column
[0176] 5) Set demand to upper and lower sub-networks
[0177] 6) Inform outmost layer about forfeiting input demands
[0178] Steps 1-5 are performed on all layers recursively followed by step 6. Steps 2, 3 and 6 are unique to the algorithm, where the other steps may be used by at least some parallel algorithms. The recursive nature of the algorithm allows the use of the same circuit for every sub- layer with a half the number of inputs–outputs per sub-network.
[0179] To implement the algorithm each PE maintains the following variables, all are reset prior to enabling the algorithm:
[0180] CollUp: Collision detected in upper sub-network
[0181] CollDown: Collision detected in lower sub-network
[0182] FirstState: State of SE in first column
[0183] LastState: State of SE in last column
[0184] Flip: SE state was flipped in previous iteration
[0185] The collision bus is denoted by CB, the output bus of each PE by OB, and the forfeit bus by FB.Attorney Ref. FRAMO-P036-WO
[0186] Summary of Heuristics in the algorithm
[0187] The following are heuristics used in the algorithm:
[0188] 1) Improved initial state
[0189] 2) Iteration sequence: publish up→publish down→flip: Get collisions from both networks then decide to flip.
[0190] 3) Flip decision: PEs with two collisions flip their state, PEs with one collision flip their state with probability 1 / 2
[0191] 4) Prev. flip state: PEs that flip their state in previous iteration do not flip their state in current iteration unless they detect two collisions
[0192] Step 1: Set demand and initial state
[0193] In this step, demands are distributed to each PE and each PE sets its initial state.
[0194] Step 2: Iterative collision detection and resolution process
[0195] An iteration has nine phases, shown in Alg.1:
[0196] Example, given the input demand in Eq. (1) and initial states of PEs in the firstcolumn is bar, that is: |{^^^^0, ^^^^1, ^^^^2, ^^^^3} = {0,0,0,0} = phase 2 operations, shown in Figure 5,are:
[0197] 1) ^^^^0: Out00= 2, Out01= 4 ^^^^1: Out10= x, Out11= 5, ...
[0198] 2) Set1={0}, Set2={2,3}, Set3={0,2}, Set4={1,3}
[0199] 3) ^^^^0sets CB[1] ← 1 since its output port is 2 (2 / / 2)
[0200] 4) ^^^^2listen to CB[1] (input 4 requires output 3) and sets CollUp ← 1 because CB[1] = 1
[0201] 5) ^^^^0sets CB[2] ← 1 since its output port is 4 (4 / / 2) ^^^^3sets CB[0] ← 1 since its output port is 0 (0 / / 2)
[0202] 6) ^^^^1listen to CB[2] (input 5 requires output 5) and sets CollDown ← 1 because CB[2] = 1
[0203] 7) ^^^^2has CollUp=1, it decides not to flip its own state ^^^^1has CollDown=1 and it decides to flip its own state
[0204] 8) ^^^^2sets CB[1] ← 1
[0205] 9) ^^^^0listen to CB[0] which is set, it flips its own stateAttorney Ref. FRAMO-P036-WO
[0206] As a result of this iteration, the states of PEs in the first column are changed to{^^^^0, ^^^^1, ^^^^2, ^^^^3} = {1,1,0,0}. This setting resolves the collision between input ports 0 and 4but input ports 1 and 3 are now colliding in the upper sub-network.
[0207] Figure 5 shows an example 500 for step 2, (a) upper sub-network (b) lower sub- network, (c) flip and flip message according to an example of the instant disclosure.
[0208] Step 3: Forfeit colliding demands
[0209] This step is to perform another iteration, as step 2, where a PE forfeits its colliding input instead of flipping its state. At the end of this step half of the colliding input ports forfeit their demands, hence reduced utilization. The randomness of selecting input demands to be forfeited guarantees fairness between demands over multiple input permutations.
[0210] 1) Phases 1-6: Same as Step 2
[0211] 2) Phase 7: PEs in sets 2 and 4 forfeit their colliding demand with probability 1 / 2.
[0212] 3) Phase 8: Same as Step 2, Publish by PEs of sets 2 and 4 that decide not to forfeit their input demand
[0213] 4) Phase 9: Same as Step 2, Listen by PEs of sets 1, 3 if accepting flip message, forfeit corresponding demand
[0214] Algorithm 1 Step 2: Iterations
[0215] Phase 2.1: Set PEs outputs based on their state
[0216] 1: Set outputs ^^^^^^,^^0, ^^^^^^,^^1based on ^^^^^^. ^^^^^^^^^^^^^^
[0217] Phase 2.2: Divide all PEs into 4 sets
[0218] 1: Set 1: PEs with valid even demand to upper sub-network
[0219] 2: Set 2: PEs with valid odd demand to upper sub-network
[0220] 3: Set 3: PEs with valid even demand to lower sub-network
[0221] 4: Set 4: PEs with valid odd demand to lower sub-network
[0222] Phase 2.3: Publish even demands to upper sub-network
[0223] 1: for ^^^^^^| {k ∈ Set1} do ^^^^^^: OB[^^^^^^,^^0 / / 2] ← 1
[0224] Phase 2.4: Detect collision in upper sub-network
[0225] 1: for each ^^^^^^| {k ∈ Set2} do
[0226] 2: if CB[^^^^^^,^^0 / / 2] = 1 then ^^^^^^. ^^^^^^^^^^^^← 1
[0227] Phase 2.5: Publish even demands to lower sub-network
[0228] 1: for ^^^^^^| {k ∈ Set3}Attorney Ref. FRAMO-P036-WO
[0229] 2: do ^^^^^^: OB[^^^^^^,^^1 / / 2] ← 1
[0230] Phase 2.6: Detect collision in lower sub-network
[0231] 1: for each ^^^^^^| {k ∈ Set4} do
[0232] 2: if CB[^^^^^^,^^1 / / 2] = 1 then ^^^^^^. ^^^^^^^^^^^^^^^^← 1
[0233] Phase 2.7: Collision resolution by listening PEs
[0234] 1: for each ^^^^^^| {k ∈ Set2 ∪ Set4} do
[0235] 2: if ^^^^^^. ^^^^^^^^^^^^= 1 and ^^^^^^. ^^^^^^^^^^^^^^^^= 1 then
[0236] 3: Flip ^^^^^^. ^^^^^^^^^^^^^^^^^^^^, ^^^^^^. ^^^^^^^^← 1
[0237] 4: else if ^^^^^^. ^^^^^^^^= 0 and
[0238] 5: (^^^^^^. ^^^^^^^^^^^^= 1 or ^^^^^^. ^^^^^^^^^^^^^^^^= 1) then
[0244] 11: end if
[0245] 12: end for
[0246] Phase 2.8: Publish flip for upper / lower sub-network
[0247] 1: for each ^^^^^^| {k ∈ Set2} do
[0248] 2: if ^^^^^^. ^^^^^^^^^^^^= 1 then
[0249] 3: ^^^^^^: OB[^^^^^^,^^0 / / 2] ← 1, ^^^^^^. ^^^^^^^^^^^^← 0
[0250] 4: for each ^^^^^^| {k ∈ Set4} do
[0251] 5: if ^^^^^^. ^^^^^^^^^^^^^^^^= 1 then
[0252] 6: ^^^^^^: OB[^^^^^^,^^1 / / 2] ← 1, ^^^^^^. ^^^^^^^^^^^^^^^^← 0
[0253] Phase 2.9: Respond to flip for upper / lower sub-network
[0254] 1: for each PEk | {k ∈ Set1 ∪ Set3} do
[0255] 2: if CB[^^^^^^,^^0 / / 2] = 1 or CB[^^^^^^,^^1 / / 2] = 1 then
[0256] 3: Flip ^^^^^^. ^^^^^^^^^^^^^^^^^^^^, ^^^^^^. ^^^^^^^^← 1
[0257] Step 4: Set states of SEs in last columnAttorney Ref. FRAMO-P036-WO
[0258] The last columns SEs’ states are set based on proposition 3 and lemma 4 using the collision bus, shown in Alg.2.
[0259] Algorithm 2 Step 4: Setting states to SEs in last column
[0260] Phase 5.1: Set states of SEs in last col. to bar, Set outputs of first col. SEs based on input state and create 4 sets
[0261] 1: for k ← 0 to N / 2 − 1 do
[0262] 2: ^^^^^^. ^^^^^^^^^^^^^^^^^^← Bar
[0263] 3: Set outputs ^^^^^^,^^0, ^^^^^^,^^1based on ^^^^^^. ^^^^^^^^^^^^^^
[0264] 4: Set1 / 2: Even / Odd demands to upper sub-network
[0265] 5: Set3 / 4: Even / Odd demands to lower sub-network
[0266] Phase 5.2: Publish demands, odd to upper even to lower
[0267] 1: for each ^^^^^^| {k ∈ Set2}
[0268] 2: ^^^^^^: OB[^^^^^^,^^0 / / 2] ← 1
[0269] 3: for each ^^^^^^| {k ∈ Set3}
[0270] 4: ^^^^^^: OB[^^^^^^,^^1 / / 2] ← 1
[0271] Phase 5.3: Set state of SEs in last column to Cross
[0272] 1: for k ← 0 to N / 2 − 1 do
[0273] 2: if CB[k] = 1 then ^^^^^^. ^^^^^^^^^^^^^^^^^^← Cross
[0274] Step 5: Deliver demands to upper / lower sub-networks
[0275] The delivery of demands from the first column to the subnetworks is via a dedicated interface identical to the natural connectivity of Beneš network.
[0276] Upper net ^^^^^^,^^0, ^^^^^^,^^1← ^^^^2^^,^^0, ^^^^2^^+1,^^0
[0277] Lower net ^^^^^^,^^0, ^^^^^^,^^1← ^^^^2^^,^^1, ^^^^2^^+1,^^1
[0278] Step 6: Inform forfeited demands
[0279] Informing the outmost layer about input demands that were forfeited is done using the dedicated forfeit bus, as described in section III-C. This process is shown in Alg.3.
[0280] Algorithm 3 Step 6: Inform outer layer of forfeited demands
[0281] phase 6.1: Innermost layer PEs publish invalid demands
[0282] 1: for k ← 0 to N / 2 − 1 do ▷ of all innermost layers
[0283] 2: if ^^^^^^,^^0is invalid then ^^^^^^: FB[^^^^^^,^^0] ← 1Attorney Ref. FRAMO-P036-WO
[0284] 3: if ^^^^^^,^^1is invalid then ^^^^^^: FB[^^^^^^,^^1] ← 1
[0285] 4: end for
[0286] phase 6.2: PEs of outmost layer listen to forfeit bus and invalidate input demands
[0287] 1: for k ← 0 to N / 2 − 1 do ▷ of all outmost layers
[0288] 2: if FB[^^^^^^,^^0] = 1 then PEk invalidates its input ^^0
[0289] 3: if FB[^^^^^^,^^1] = 1 then PEk invalidates its input ^^1
[0290] 4: end for
[0291] Termination of recursion
[0292] The recursive iterative process stops at 8 × 8 sub-networks which are resolved inO(1) time in 2 steps as follows:
[0293] 1) Resolve first and last columns in 8Å~8 Beneš network:
[0294] The state of one SE in the first column can be fixed to cross or bar and complete routing can still be found. Once fixing a state to an SE, three SEs remain unset, resulting in 8 (23) setting options. A simple combinatorial logic evaluates these 8 options in parallel within O(1) time.
[0295] 2) Resolve entire 4 × 4 Beneš network:
[0296] A pre-configured 24-entry (4! Input permutations) table is used. Each entry contains the states of all six SEs of the 4x4 network. Partial permutations are filled prior to accessing the table. Permutation filling and table access are both within O(1) time.
[0297] IMPLEMENTATION AND TIMING BUDGET
[0298] Collision bus sizing and related pipeline
[0299] A single clock cycle may not be sufficient to deliver and process messages on the collision bus, especially for highradix networks. Therefore, a pipeline design can be used such that one clock cycle is dedicated to propagate a message and sampled at the receiving PEs and another clock cycle to process the received message. Figure 6 shows a single iteration pipeline 600 according to an example of the instant disclosure.
[0300] N / 2 bits collision bus: 6 clock cycles per iteration. Cycles 0, 1 are used to deliver collision messages for upper and lower sub-networks. The messages are accepted in cycles 1 and 2. Cycle 4 is used to deliver the flip messages as shown in Figure 6.Attorney Ref. FRAMO-P036-WO
[0301] N bit collision bus: 4 clock cycles per iteration. Cycle 0 is used to publish demands for both upper and lower networks and cycle 2 is used to send flip messages as shown in Figure 6.
[0302] Relating number of iterations to clock cycles
[0303] Given I as the number of iterations per layer, the number of clock cycles per stepfor N / 2 bits Collision bus are: step 2: 6 ∙ ^^ clocks, step 3: 6 clocks, steps 4 and: 3 clocks each.Resulting in I +2 Iterations which are 6 ∙ (^^ +2) clock cycles.
[0304] For N bits Collision bus, step 2 takes 4 ∙ ^^ clocks, step 3: 4 clocks, steps 4 and 6: 3clocks each. Resulting in I + 2.5 Iterations which are 4 ∙ (^^ + 25) clock cycles.
[0305] Step 5 (delivery of demands to the two sub-networks), is implementation dependent, and therefore is not shown.
[0306] In order to complete the routing of the entire NxN network within ^^((^^^^^^2^^)2) iterations it is possible to limit I which is derived by comparing the number of iterations for all layers (^^^^^^2(N / 8) layers), shown above, to (^^^^^^2^^)2for N bits collision bus as depicted in Eq.10.
[0307] (log22^^)= (^^ + 2.5) ∙ log2(^^ / 8) ^ ^^ = (log2 ^^)2 / log2(^^ / 8) − 2.5 (10)
[0308] 12.
[0309] The total number of clock cycles is derived by multiplying the total number of iterations by 4 or 6 (depending on the size of the collision bus). e.g. N=32 and I=10 yields 100 cycles.
[0310] Summary of busses and sizing
[0311] Routing the entire NÅ~N Benesˇ network requires only N / 2 PEs. The outmost layer uses all N / 2 PEs along with the entire collision bus and the entire forfeit bus. Next layer contains two sub-networks where each requires N / 4 PEs along with half of the collision bus and half of the forfeit bus and so on.
[0312] A summary of interfaces and connectivity per PE is shown in Table I. Reduced or extended collisions bus are considered based on implementation constraints. Other implementation dependent interfaces are not shown.
[0313] TABLE I: Summary of required wiring Per PE For all PEs Reduce Coll. bus N / 2 in, N / 2 out (N)^^2 / 2Extend Coll. bus N in, N out (2N)^^2Attorney Ref. FRAMO-P036-WO Forfeit bus N in, N out (2N) ^^2Dynamic Iters. 1 in, 1 out (2) N Alg. Step 6 1 in, 1 out (2) N
[0314] EVALUATION
[0315] It is possible to evaluate the performance of the routing algorithm for Benesˇ networks as a function of network size (N=32...1024) and number of iterations. The initial state of PEs is as defined in algorithm step 1.
[0316] As mentioned herein, Benesˇ networks are currently mainly suggested as an attractive solution for optical cross-connects. In this context, the full features of EPSs such as queuing engine, congestion control, scheduling, and more, that greatly affect the input permutations patterns to the cross connect, are not relevant to our proposed algorithm. Therefore random input permutations are used (full and partial) with no specific workload related input. Furthermore, any traffic patterns such as many-to-1 cause a blocking effect which are not relevant for the algorithm as they potentially create partial permutations.
[0317] First column utilization
[0318] Average utilization of the first column only was measured for 100% and 90% input loads, shown in Figure 7. Each point is calculated based on 50,000 random permutations. High average utilization, >95%, is achieved with just few iterations.
[0319] Figure 7 shows first column average utilization 700 according to an example of the instant disclosure.
[0320] The utilization distribution of the first column only is shown in Figure 8 for 100% and 90% input load.
[0321] At 100% load, the average number of iterations is 10.0 for N=32, 25.4 for N=64, for 82.6 for N=128 and 131.4 for N=256. At 90% load, 8.6 for N=32, 15.8 for N=64, 27.5 for N=128 and 46.4 for N=256. High number of iterations is required to achieve 100% utilization even for relatively small networks. The long tails are due to rare cyclic behavior of the collision resolution algorithm that may randomly return to a PE state already visited. As 90% input load imposes less constraints on the routing algorithm, therefore, higher utilization is achieved with less iterations. For example, N=256 at 90% requires a third of the number required for 100% load.Attorney Ref. FRAMO-P036-WO
[0322] Figure 8 shows a histogram of the number of iterations 800 according to an example of the instant disclosure.
[0323] As it is desirable to want the algorithm to perform within ^^(log2^^)2steps (iterations), best known time complexity for parallel algorithms, it is possible to limit the total number of iterations as shown in Eq.10, achieving less than 100% utilization but with less connectivity between PEs which allows for easier scalability to high-radix OCSs.
[0324] Fixed number of iterations per layer
[0325] The disclosure examines the algorithm performance for the entire Beneš networkwhere the number of iterations (I) is fixed per layer where layers 8 × 8 and 4 × 4 are resolved inO (1). In total I ∙ log2 ^^ / 8 iterations are performed in the entire network. The initial state of PEsis as described in algorithm step 1. Results are shown in Figure 9.
[0326] Figure 9 shows average utilization of ^^ × ^^ network 900, fixed mode for (a)100% input load and for (b) 90% input load according to an example of the instant disclosure.
[0327] At 100% input load the algorithm achieves above 95% average utilization for 32× 32 with I≥6 and for 64× 64 with I≥10 and higher network sizes achieve above 92% averageutilization all using at most (log2^^)2total iterations.
[0328] Figure 10 shows distribution of Utilization for 100% input demands for 12 iterations per layer 1000 according to an example of the instant disclosure.
[0329] The disclosure describes a novel approach to Beneš network routing algorithm and architecture by using an iterative algorithm that trades-off between utilization and completion time. The algorithm is scalable to high-radix networks suitable for OCSs, achieved by a reduced connectivity complexity.
[0330] The reduction in the connectivity complexity and the comparable time complexity is achieved at the cost of having average utilization below but nearly 100%, using fairly small amount of logic per each PE without the usage of special blocks such as SRAM or DSP.
[0331] The algorithm achieves >94% utilization for full input demands on network sizesup to 128 × 128 and >92% utilization on network sizes up to 1024 × 1024. The algorithmachieves even higher output utilization for partial permutations. In fact, it performs better for partial permutations.
[0332] APPENDIX A: PROOFS OF LEMMAS
[0333] Proof to lemma 1:Attorney Ref. FRAMO-P036-WO
[0334] Proof. Since a SE in the last column is connected to 2 output ports, it is not possible for that SE to be requested by another input demand (input demand = output port).
[0335] Proof to lemma 2:
[0336] Proof. Flipping a state of a colliding SE in the first column resolve the collision to the corresponding SE in the last column as the colliding demand is now routed through the other subnetwork. However, flipping the state of the colliding SE brings another demand to the sub- network which may collide with another demand already routed in this sub-network.
[0337] Proof to lemma 3:
[0338] Proof. Similar to lemma 2 proof extended to two collisions.
[0339] Proof to lemma 4:
[0340] Proof. No collisions from the first column means traffic to ^^^^^^in the last column arrives one from the upper sub-network and one from the lower sub-network. This means even and odd outputs for ^^^^^^in the last column can not arrive from the same sub-network and therefore collisions are not possible.
[0341] Proof to lemma 5:
[0342] Proof. A collision between two input demands to ^^^^^^in the last column through one sub-network implies that there is no demand to ^^^^^^from the other sub-network. Therefore, the N / 2 input demands connected to this other sub-network has only N / 2−1 available destination ports, using the pigeonhole principle, two requests must share one of the output ports of the other sub-network, hence a collision in the other subnetwork as well.
[0343] Figure 11 shows a method of routing a Beneš network 1100 according to an example of the instant disclosure. Although the example method 1100 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the method 1100. In other examples, different components of an example device or system that implements the method 1100 may perform functions at substantially the same time or in a specific sequence.
[0344] According to some examples, the method 1100 includes for a plurality of processing elements that represent a plurality of switching elements in a Beneš network and a communication bus in between the plurality of processing elements, wherein two sub-networks are disposed in between a first column of the Beneš network and a last column of the BenešAttorney Ref. FRAMO-P036-WO network, wherein the first column comprises a first set of switching elements in the plurality of switching elements, and wherein the last column comprises a second set of switching elements in the plurality of switching elements and setting an initial state for the plurality of processing elements that represent switching elements of the first column at block 1110.
[0345] According to some examples, the method 1100 includes configuring the plurality of processing elements that represent the switching elements of the first column to communicate in a communication with each other to detect a plurality of collisions and to resolve the plurality of collisions by flipping states in a plurality of iterations, the communication performed by a centralized collision bus at block 1120.
[0346] According to some examples, the method 1100 includes when the plurality of iterations are completed, resolving any remaining collisions, if any, by forfeiting an input demand from one colliding input port per a remaining collision, resulting in a reduction of utilization at block 1130.
[0347] In an example, each bit in the centralized collision bus represents a switching element in the second set of switching elements.
[0348] In another example, the detection of the plurality of collisions is performed by the centralized collision bus, the centralized collision bus provides input to the plurality of processing elements, each bit in the centralized collision bus represents a switching element in the last column of the Beneš network, and the method 1100 further comprises, for each iteration in the plurality of iterations the method 1100 may further include setting, by a set of publisher elements in the plurality of processing elements, bits on the centralized collision bus and observing, by a set of listener elements in the plurality of processing elements, the set of listener elements opposite the set of publisher elements, corresponding bits on the centralized collision bus, to detect a collision in the plurality of collisions when the bits are set.
[0349] In another example, in method 1100 the set of publisher elements may include processing elements having demands to a first set of output ports that are routed through a first sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to a second set of output ports that are routed through the first sub- network. Additionally, in method 1100 the set of publisher elements may include processing elements having demands to the second set of output ports that are routed through a first sub-Attorney Ref. FRAMO-P036-WO network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the first sub-network.
[0350] Additionally, in method 1100, the set of publisher elements may include processing elements having demands to the second set of output ports that are routed through a second sub- network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the second sub-network.
[0351] Additionally, in method 1100, the set of publisher elements may include processing elements having demands to the first set of output ports that are routed through a second sub- network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the second set of output ports that are routed through the second sub-network.
[0352] In another example, the first set of output ports are even output ports and wherein the second set of output ports are odd output ports.
[0353] In another example, the flipping states may include when a listener element in the set of listener elements detects two collisions in the plurality of collisions, the listener element flips its state and does not send a flip state message, when the listener element detects one collision in the plurality of collisions, the listener element flips its state randomly, with a probability of 0.5, and does not send the flip state message, and when the listener element sends the flip state message, the listener element does not change its state and a corresponding publisher element in the set of publisher elements flips its state upon detecting the flip state message, thereby setting a bit corresponding to a switching element in the last column.
[0354] In another example, setting the initial state further includes directing all input demands to a first set of output ports to a first sub-network in the two sub-networks and directing all input demands to a second set of output ports to a second sub-network in the two sub-networks or setting the initial state as a random state when a given processing element in the plurality of processing elements has demands to two of the first set of output ports or two of the second set of output ports.
[0355] Figure 12 shows another method of routing a Beneš network 1200 according to an example of the instant disclosure. Although the example method 1200 depicts a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the method 1200. In otherAttorney Ref. FRAMO-P036-WO examples, different components of an example device or system that implements the method 1200 may perform functions at substantially the same time or in a specific sequence.
[0356] According to some examples, the method 1200 includes distributing a plurality of demands to a plurality of processing elements in the Beneš network at block 1210.
[0357] According to some examples, the method 1200 includes setting an initial state for each of the plurality of processing elements at block 1220.
[0358] According to some examples, the method 1200 includes performing an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network at block 1230.
[0359] According to some examples, the method 1200 includes forfeiting, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision at block 1240.
[0360] According to some examples, the method 1200 includes setting states of a plurality of switching elements in a last column of the Beneš network at block 1250.
[0361] According to some examples, the method 1200 includes delivering at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network at block 1260.
[0362] According to some examples, the method 1200 includes either (i) informing, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) informing an outermost layer of the Beneš network of zero forfeited and colliding demands at block 1270. In an example, the method 1200 is performed recursively and repeatedly for each of the one or more sub-networks in parallel, ending in the smallest sub-network of the one or more sub-networks.
[0363] According to an example, the setting is performed by a centralized collision bus, and the informing is performed by a dedicated forfeit bus.
[0364] According to an example, the centralized collision bus is one of: (i) a collision bus with N / 2 bits and at least one clock cycle per iteration for delivering collision messages, accepting the collision messages, and delivering flip state messages, and (ii) a collision bus with N bits and at least one clock cycle per iteration for delivering collision messages and flipping state messages.Attorney Ref. FRAMO-P036-WO
[0365] According to an example, the recursive and repeated performance of the method 1200 stops at 8x8 sub-networks of the Beneš network, and the method 1200 resolves within a time of O(1) when the Beneš network has the 8x8 sub-networks or fewer.
[0366] Figure 13 shows an example of computing system 1300, which can be, for example, a computing device, or any component thereof in which the components of the system are in communication with each other using connection 1305. Connection 1305 can be a physical connection via a bus, or a direct connection into processor 1310, such as in a chipset architecture. Connection 1305 can also be a virtual connection, networked connection, or logical connection.
[0367] In some embodiments, computing system 1300 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple data centers, a peer network, etc. In some embodiments, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some embodiments, the components can be physical or virtual devices.
[0368] Example system 1300 includes at least one processing unit (CPU or processor) 1310 and connection 1305 that couples various system components including system memory 1315, such as read-only memory (ROM) 1320 and random access memory (RAM) 1325 to processor 1310. Computing system 1300 can include a cache of high-speed memory 1312 connected directly with, in close proximity to, or integrated as part of processor 1310.
[0369] Processor 1310 can include any general purpose processor and a hardware service or software service, such as services 1332, 1334, and 1336 stored in storage device 1330, configured to control processor 1310 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1310 may essentially be a completely self-contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.
[0370] To enable user interaction, computing system 1300 includes an input device 1345, which can represent any number of input mechanisms, such as a microphone for speech, a touch- sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1300 can also include output device 1335, which can be one or more of a number of output mechanisms known to those of skill in the art. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1300. Computing system 1300 can include communications interface 1340,Attorney Ref. FRAMO-P036-WO which can generally govern and manage the user input and system output. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0371] Storage device 1330 can be a non-volatile memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, random access memories (RAMs), read-only memory (ROM), and / or some combination of these devices.
[0372] The storage device 1330 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1310, it causes the system to perform a function. In some embodiments, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1310, connection 1305, output device 1335, etc., to carry out the function.
[0373] For clarity of explanation, in some instances, the present technology may be presented as including individual functional blocks including functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software.
[0374] Any of the steps, operations, functions, or processes described herein may be performed or implemented by a combination of hardware and software services or services, alone or in combination with other devices. In some embodiments, a service can be software that resides in memory of a client device and / or one or more servers of a content management system and perform one or more functions when a processor executes the software associated with the service. In some embodiments, a service is a program or a collection of programs that carry out a specific function. In some embodiments, a service can be considered a server. The memory can be a non- transitory computer-readable medium.
[0375] In some embodiments, the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.Attorney Ref. FRAMO-P036-WO
[0376] Methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer-readable media. Such instructions can comprise, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or special purpose processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The executable computer instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, solid-state memory devices, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0377] Devices implementing methods according to these disclosures can comprise hardware, firmware and / or software, and can take any of a variety of form factors. Typical examples of such form factors include servers, laptops, smartphones, small form factor personal computers, personal digital assistants, and so on. The functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0378] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are means for providing the functions described in these disclosures.
Claims
Attorney Ref. FRAMO-P036-WO CLAIMS What is claimed is:
1. A method for routing a Beneš network, the method comprising: for a plurality of processing elements that represent a plurality of switching elements in a Beneš network and a communication bus in between the plurality of processing elements, wherein two sub-networks are disposed in between a first column of the Beneš network and a last column of the Beneš network, wherein the first column comprises a first set of switching elements in the plurality of switching elements, and wherein the last column comprises a second set of switching elements in the plurality of switching elements: setting an initial state for the plurality of processing elements that represent switching elements of the first column; configuring the plurality of processing elements that represent the switching elements of the first column to communicate in a communication with each other to detect a plurality of collisions and to resolve the plurality of collisions by flipping states in a plurality of iterations, the communication performed by a centralized collision bus; and when the plurality of iterations are completed, resolving any remaining collisions, if any, by forfeiting an input demand from one colliding input port per a remaining collision, resulting in a reduction of utilization, wherein each bit in the centralized collision bus represents a switching element in the second set of switching elements.
2. The method of claim 1, wherein the detection of the plurality of collisions is performed by the centralized collision bus, the centralized collision bus provides input to the plurality of processing elements, each bit in the centralized collision bus represents a switching element in the last column of the Beneš network, and the method further comprises, for each iteration in the plurality of iterations: setting, by a set of publisher elements in the plurality of processing elements, bits on the centralized collision bus; and observing, by a set of listener elements in the plurality of processing elements, the set of listener elements opposite the set of publisher elements, corresponding bits on the centralized collision bus, to detect a collision in the plurality of collisions when the bits are set.Attorney Ref. FRAMO-P036-WO 3. The method of claim 2, wherein at least one of: (i) the set of publisher elements comprises processing elements having demands to a first set of output ports that are routed through a first sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to a second set of output ports that are routed through the first sub-network, (ii) the set of publisher elements comprises processing elements having demands to the second set of output ports that are routed through a first sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the first sub-network, (iii) the set of publisher elements comprises processing elements having demands to the second set of output ports that are routed through a second sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the second sub-network, and (iv) the set of publisher elements comprises processing elements having demands to the first set of output ports that are routed through a second sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the second set of output ports that are routed through the second sub-network.
4. The method of claim 3, wherein the first set of output ports are even output ports and wherein the second set of output ports are odd output ports.
5. The method of claim 2, wherein the flipping states further comprises: when a listener element in the set of listener elements detects two collisions in the plurality of collisions, the listener element flips its state and does not send a flip state message; when the listener element detects one collision in the plurality of collisions, the listener element flips its state randomly, with a probability of 0.5, and does not send the flip state message; when the listener element sends the flip state message, the listener element does not change its state and a corresponding publisher element in the set of publisher elements flips its state upon detecting the flip state message, thereby setting a bit corresponding to a switchingAttorney Ref. FRAMO-P036-WO element in the last column.
6. The method of claim 1, wherein the setting the initial state further comprises one of : directing all input demands to a first set of output ports to a first sub-network in the two sub-networks and directing all input demands to a second set of output ports to a second sub- network in the two sub-networks; and setting the initial state as a random state when a given processing element in the plurality of processing elements has demands to two of the first set of output ports or two of the second set of output ports.
7. A method for routing a Beneš network, the method comprising: distributing a plurality of demands to a plurality of processing elements in the Beneš network; setting an initial state for each of the plurality of processing elements; performing an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network; forfeiting, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision; setting states of a plurality of switching elements in a last column of the Beneš network; delivering at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network; and either (i) informing, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) informing an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the method is performed recursively and repeatedly for each of the one or more sub-networks in parallel, ending in the smallest sub-network of the one or more sub-networks.
8. The method of claim 7, wherein the setting is performed by a centralized collision bus, and wherein the informing is performed by a dedicated forfeit bus.Attorney Ref. FRAMO-P036-WO 9. The method of claim 8, wherein the centralized collision bus is one of: (i) a collision bus with N / 2 bits and at least one clock cycle per iteration for delivering collision messages, accepting the collision messages, and delivering flip state messages, and (ii) a collision bus with N bits and at least one clock cycle per iteration for delivering collision messages and flipping state messages.
10. The method of claim 7, wherein the recursive and repeated performance of the method stops at 8x8 sub-networks of the Beneš network, and wherein the method resolves within a time of O(1) when the Beneš network has the 8x8 sub-networks or fewer.
11. A system for routing a Beneš network, the system comprising: a plurality of processing elements that represent a plurality of switching elements in a Beneš network and a communication bus in between the plurality of processing elements, wherein two sub-networks are disposed in between a first column of the Beneš network and a last column of the Beneš network, wherein the first column comprises a first set of switching elements in the plurality of switching elements, and wherein the last column comprises a second set of switching elements in the plurality of switching elements, the plurality of processing elements: setting an initial state for the plurality of processing elements that represent switching elements of the first column; configuring the plurality of processing elements that represent the switching elements of the first column to communicate in a communication with each other to detect a plurality of collisions and to resolve the plurality of collisions by flipping states in a plurality of iterations, the communication performed by a centralized collision bus; and when the plurality of iterations are completed, resolving any remaining collisions, if any, by forfeiting an input demand from one colliding input port per a remaining collision, resulting in a reduction of utilization, wherein each bit in the centralized collision bus represents a switching element in the second set of switching elements.
12. The system of claim 11, wherein the detection of the plurality of collisions isAttorney Ref. FRAMO-P036-WO performed by the centralized collision bus, the centralized collision bus provides input to the plurality of processing elements, each bit in the centralized collision bus represents a switching element in the last column of the Beneš network, and for each iteration in the plurality of iterations: setting, by a set of publisher elements in the plurality of processing elements, bits on the centralized collision bus; and observing, by a set of listener elements in the plurality of processing elements, the set of listener elements opposite the set of publisher elements, corresponding bits on the centralized collision bus, to detect a collision in the plurality of collisions when the bits are set.
13. The system of claim 12, wherein at least one of: (i) the set of publisher elements comprises processing elements having demands to a first set of output ports that are routed through a first sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to a second set of output ports that are routed through the first sub-network, (ii) the set of publisher elements comprises processing elements having demands to the second set of output ports that are routed through a first sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the first sub-network, (iii) the set of publisher elements comprises processing elements having demands to the second set of output ports that are routed through a second sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the first set of output ports that are routed through the second sub-network, and (iv) the set of publisher elements comprises processing elements having demands to the first set of output ports that are routed through a second sub-network in the two sub-networks, and the set of listener elements comprises processing elements having demands to the second set of output ports that are routed through the second sub-network.
14. The system of claim 13, wherein the first set of output ports are even output ports and wherein the second set of output ports are odd output ports.Attorney Ref. FRAMO-P036-WO 15. The system of claim 12, wherein the flipping states further comprises: when a listener element in the set of listener elements detects two collisions in the plurality of collisions, the listener element flips its state and does not send a flip state message; when the listener element detects one collision in the plurality of collisions, the listener element flips its state randomly, with a probability of 0.5, and does not send the flip state message; and when the listener element sends the flip state message, the listener element does not change its state and a corresponding publisher element in the set of publisher elements flips its state upon detecting the flip state message, thereby setting a bit corresponding to a switching element in the last column.
16. The system of claim 11, wherein the setting the initial state further comprises one of: directing all input demands to a first set of output ports to a first sub-network in the two sub-networks and directing all input demands to a second set of output ports to a second sub- network in the two sub-networks; and setting the initial state as a random state when a given processing element in the plurality of processing elements has demands to two of the first set of output ports or two of the second set of output ports.
17. A system for routing a Beneš network, the system comprising: a plurality of processing elements to: receive a plurality of demands in the Beneš network; set an initial state for each of the plurality of processing elements; perform an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network; forfeit, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision; set states of a plurality of switching elements in a last column of the Beneš network; deliver at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network; andAttorney Ref. FRAMO-P036-WO either (i) inform, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) inform an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the system is a recursive system to repeatedly for each of the one or more sub- networks in parallel, end in the smallest sub-network of the one or more sub-networks using recursion.
18. The system of claim 17, wherein the setting of the initial state is performed by a centralized collision bus, and wherein the informing is performed by a dedicated forfeit bus.
19. The system of claim 18, wherein the centralized collision bus is one of: (i) a collision bus with N / 2 bits and at least one clock cycle per iteration for delivering collision messages, accepting the collision messages, and delivering flip state messages, and (ii) a collision bus with N bits and at least one clock cycle per iteration for delivering collision messages and flipping state messages.
20. The system of claim 17, wherein the recursion stops at 8x8 sub-networks of the Beneš network, and resolves within a time of O(1) when the Beneš network has the 8x8 sub- networks or fewer.
21. A non-transitory computer-readable storage medium, having instructions stored thereon that, when executed by at least one computing device cause the at least one computing device to perform operations comprising: distributing a plurality of demands to a plurality of processing elements in the Beneš network; setting an initial state for each of the plurality of processing elements; performing an iterative collision detection and resolution process to set states of a plurality of switching elements in a first column of the Beneš network; forfeiting, by at least one processing element in the plurality of processing elements, at least one colliding input, to prevent at least one collision; setting states of a plurality of switching elements in a last column of the Beneš network;Attorney Ref. FRAMO-P036-WO delivering at least one demand from the first column of the Beneš network to one or more sub-networks of the Beneš network; and either (i) informing, at an end of a recursion, an outermost layer of the Beneš network of the at least one colliding input that was forfeited, or (ii) informing an outermost layer of the Beneš network of zero forfeited and colliding demands, wherein the method is performed recursively and repeatedly for each of the one or more sub-networks in parallel, ending in the smallest sub-network of the one or more sub-networks.