Bridge device for use with an optical switching network or an optical link

The bridge device with a state machine and buffer circuit addresses routing complexity in optical networks, enabling efficient, low-power data transmission with reduced reconfiguration times and minimal disruption.

WO2026068949A1PCT designated stage Publication Date: 2026-04-02SALIENCE LABS LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2026-04-02

AI Technical Summary

Technical Problem

Optical switching networks face complexity in routing and head-of-line blocking as networks grow large, requiring buffers and opto-electric conversions that increase complexity and power consumption.

Method used

A bridge device with a state machine for routing data streams, incorporating a buffer circuit and bypass path on the data plane, and a control plane for managing data transmission and buffer states, allowing optical signal reconditioning without opto-electric conversions.

Benefits of technology

Enables efficient, low-power data routing in large optical switching networks with reduced reconfiguration times and minimal disruption, maintaining optical signal integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure GB2025052098_02042026_PF_FP_ABST
    Figure GB2025052098_02042026_PF_FP_ABST
Patent Text Reader

Abstract

A bridge device (500) for use with an optical switching network or an optical link is provided. The bridge device (500) includes: a first port circuit (510) for receiving a data stream, a second port circuit (520) for transmitting the data stream, and a buffer circuit (530) and / or bypass path provided on a data plane. The bridge device (500) may also include a state machine (540) provided on a control plane. The state machine receives routing information for routing the data stream between a source host and a destination host in the optical switching network or optical link and instructs the buffer circuit (530) to either store the received data stream or to pass the data stream to the second port circuit (520), based on the routing information, and buffer state information indicating whether the buffer is full or has available memory space.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] BRIDGE DEVICE FOR USE WITH AN OPTICAL SWITCHING NETWORK OR AN OPTICAL LINK

[0002] Technical Field

[0003] The present disclosure relates to a bridge device for use with an optical switching network, or with an optical link.

[0004] Background

[0005] Optical circuit switching (OCS) permits exchange of data in an optical form between a source host and a destination host. Compared with electrical switching networks, the optical counterpart benefits from a broader bandwidth and lower power consumption. As such, they may be considered for many applications.

[0006] When a switching network becomes relatively large, the routing through the network becomes more complex. Algorithms which establish end-to-end paths through a network (for instance using "Virtual cut through routing” or "Wormhole routing”) suffer from head-of-line blocking once the networks grow large.

[0007] Buffers may be introduced at every point where a routing decision needs to be made. The buffers then allow to halt the traffic when the desired route through a switch is already in use (congestion) or to also allow for interrupting an existing path through a switch whose stream of data that it carries has lower priority. The disrupted path can then keep its data in the respective buffer and avoid retransmission all the way from source to sink through the network.

[0008] In electrical networks, this buffering is implemented in all switches and works based on packets. The typical electrical switch operates on OSI layers 3 (sometimes 4) and has intricate knowledge of the nature of data that is passing through it. When combined with optical links between such electrical switches, this approach comes at the cost of having to do an opto-electric conversion at the input to the switch as well as an electro-optic conversion at the output. In large networks, the optical signal may also need to undergo reconditioning. This usually involves opto-electric and electro-optic reconversion as the reconditioning is performed on digital data.

[0009] It is an object of the disclosure to address one or more of the above- mentioned limitations.

[0010] Summary

[0011] According to a first aspect of the disclosure, there is provided a bridge device for use with an optical switching network or an optical link, the bridge device comprising: a first port circuit for receiving a data stream; a second port circuit for transmitting the data stream; a buffer circuit and / or bypass path provided on a data plane.

[0012] Optionally, the bridge device comprises a state machine provided on a control plane; the state machine being configured to receive routing information for routing the data stream between a source host and a destination host in the optical switching network or optical link, and to instruct the buffer circuit to either store the received data stream or to pass the data stream to the second port circuit, based on the routing information, and buffer state information indicating whether the buffer is full or has available memory space.

[0013] For instance, routing information may include ID of a source host and a destination host. Optionally, the buffer circuit comprises a per source / destination pair buffer.

[0014] Optionally, wherein the first port circuit comprises a first physical medium attach coupled to both a data plane receiver and a control plane receiver; and wherein the second port circuit comprises a second physical medium attach coupled to both a data plane transmitter and a control plane transmitter.

[0015] Optionally, the state machine is configured to receive transmitter information indicating if the data plane transmitter is available for transmitting data. For instance, the transmitter information may also include information indicating if the data plane transmitter is connected to a specific destination port.

[0016] Optionally, wherein the control plane receiver is configured to receive instructions to perform one or more of: deserialize data on the data plane, recover a clock signal, process control characters.

[0017] Optionally, wherein the state machine is coupled to a first circuit and to a second circuit, wherein each one of the first circuit and the second circuit has a data plane and a control plane.

[0018] Optionally, wherein the first circuit is coupled to a first bi-directional physical medium attach; and wherein the second circuit is coupled to a second bi-directional physical medium attach.

[0019] Optionally, wherein the data plane comprises a data physical coding sublayer, coupled to a data physical sublayer.

[0020] Optionally, wherein the control plane comprises a set of buffers, a control physical coding sublayer, a control physical sublayer, a protocol circuit, and a phase locked loops and clock distribution circuit. Optionally, the bridge device comprises a translation circuit configured to perform network address translation. For instance, a host may have a different address in different domains, and the translation circuit may be configured to translate a host address in one domain to a corresponding address in another domain.

[0021] Optionally, wherein the bridge device is configured to receive network protocol information and wherein the bridge device is protocol agnostic.

[0022] Optionally, the bridge device comprises a memory.

[0023] Optionally, wherein the first port circuit is adapted to perform optical to electrical conversion, and wherein the second port circuit is adapted to perform electrical to optical conversion. For instance, the optical to electrical conversion and the electrical to optical conversion may be used to recondition an optical signal.

[0024] According to a second aspect of the disclosure, there is provided a bridge device system comprising at least one pair of bridge devices, wherein the said at least one pair of bridge devices comprises first bridge device according to the first aspect and configured to transmit data in a first direction, and a second bridge device according to the first aspect configured to transmit data in a second direction.

[0025] Optionally, for each pair, the state machine of the first bridge device is coupled to the state machine of the second bridge device. For instance, the state machine of the first bridge device and the state machine of the second bridge device may be connected to each other or unified to form a single state machine. Optionally, the first bridge device is configured to perform a primary conversion from a first optical standard on a receiving side to a second optical standard on a transmitting side and wherein the second bridge device is configured to perform conversion from the second optical standard on a receiving side to the first optical standard on a transmitting side.

[0026] For instance, the first standard may encode data using a first predetermined number of wavelengths and / or other optical signalling properties, and the second standard may encode data using a second predetermined number of wavelengths and / or other optical signalling properties. For example, the first standard may be 400G-FR4 and the second standard 400G-DR4 or the first standard may be IEEE standard and the second standard a custom optical transmission specification.

[0027] Optionally, wherein the first bridge device has a first pair of physical medium attachments for performing the primary conversion; and wherein the second bridge device has a second pair of physical medium attachments for performing the secondary conversion.

[0028] According to a third aspect of the disclosure, there is provided an optical switching network comprising a plurality of bridge devices according to the first aspect.

[0029] Optionally, the optical switching network comprises a plurality of switch devices, each switch device being associated with a corresponding domain; and wherein a bridge device or a bridge device pair is provided between two different domains.

[0030] Optionally, the optical switching network comprises a plurality of hosts, and wherein a host has a different address in different domains. Optionally, wherein each host comprises at least one network interface controller; a network data plane and a network control plane distributed among the plurality of hosts and the plurality of switch devices forming the network; wherein the said at least one network interface controller is configured to provide network protocol information to the network data plane and control information to the network control plane for routing optical data to be transmitted between a source host and a destination host among the plurality of hosts.

[0031] Description of the drawings

[0032] The disclosure is described in further detail below by way of example and with reference to the accompanying drawings, in which: figure 1A is a schematic diagram of an optical switching network according to the disclosure; figure IB shows various layers that may be implemented in the optical switching network of figure 1A; figure 2A is a diagram of a switch device for use in the optical switching network of figure 1; figure 2B is a diagram of another switch device for use in the optical switching network of figure 1; figure 3A is a partial diagram of a host NIC for use in the optical switching network of figure 1; figure 3B is a diagram of an exemplary implementation of a transceiver for use in the circuit of figure 3A; figure 3C is a diagram of an example implementation of a PLLs and clock distribution circuit for use in figure 3A; figure 4 is a diagram of an optical switching network comprising several bridge devices; figure 5 is a diagram of a bridge device for use in the optical switching network of figure 4; figure 6 is a diagram of bridge device system for use in the system of figure 4; figure 7 is an example implementation of a bridge device pair; figure 8 is an example implementation of a smartNIC data and control planes for use in the bridge device pair of figure 7 A; figure 9 is a diagram of an exemplary physical architecture of a bridge device; figure 10 is a diagram illustrating communication between two host devices using two switch devices connected to each other; figure 11 is a diagram illustrating communication between two host devices using two switch devices connected to each other via a bridge device pair.

[0033] Description

[0034] Various acronyms and abbreviations are used in the present disclosure and listed in the following glossary.

[0035] Host - a host node (source and sink of application data) attached to the network via a network interface controller (NIC).

[0036] NIC - network interface controller containing the MAC / PCS / PHYs for data and control plane.

[0037] Network protocol - The protocol (stack) implemented on the data plane. The switch devices described in the disclosure are agnostic to these protocols, however the network interface controllers (NICs) are network protocol aware, at least to some degree. In practice this means that the NIC is designed to implement a procedure for embedding or de-embedding data into the packet layer (L3) of the network protocol. PCS - Physical coding sublayer, contains protocol flit / package detection, lane to lane deskew, line code embeddings (e.g. 64 / 66 or 128 / 130b), data de- / scrambling and error correction (Forward error correction FEC) and depending on protocol, the flow control.

[0038] MAC - Medium access control sublayer. Although this is a term customarily used in Ethernet, in the context of this disclosure MAC is used to denote a layer with the following properties: i) Seen from the point of higher layers (towards OS and software) the first to potentially establish hardware-based flow control, unless the PCS takes care of this functionality. For Ethernet, this means the Reconciliation sublayer (RS) is attributed to the MAC layer. ii) Seen from the point of higher layers, the last layer to have plain and well-defined access to the destination host ID on the network.

[0039] PHY - Physical sublayer (connecting to the medium of transmission), comprising clocking (PLL), serialization, deserialization, equalization, driver and receiver amplifiers, clock data recovery circuits. Ethernet often separate PHY into physical medium attach (PMA) and physical medium dependent (PMD).

[0040] Link - established bidirectional serial data transmission over one or many physical connections (lanes) or optical channels between a specific host and a specific switch device, attached to a single port on either side. Optical channels may be provided on a same fiber or may be distributed over multiple fibers. Each optical channel may carry a signal (for instance data) at a specific wavelength A.

[0041] Lane - a physical electrical connection from one point to another. In case of differential signalling, the connection is made using two physical conductors. Channel - the optical equivalent to a lane. A wavelength (lambda) of specific polarization with associated, dedicated bandwidth around it, onto which the information can be encoded. Channels of different wavelength, polarization and modes may also be encoded onto physically separated fibers.

[0042] Port - taken to be bidirectional and including physical data and control plane connection (lanes / channels). The number of ports of a switch is called the radix.

[0043] The terms transceiver, SerDes and Serializer are used interchangeably throughout the description. They signify the union of one transmitter with data serialization path, analog signal equalization as well as (electrical and optical) driver circuitry, one receiver with data deserializer, analog frontend (TIA, equalization stages such as FFE) and the clock data recovery (CDR) circuit as well as the clock distribution circuitry required to operate all circuits. For clarity, the phase locked loop (PLL) circuits are excluded. The PLL circuits are customarily deployed in these systems to generate high frequency, low jitter clock signals for the transceiver circuit.

[0044] Other common abbreviations include: FFE - Feed forward equalization; FIR - Finite impulse response filter; FEC - Forward error correction; DFE - Decision feedback equalization; CDR - Clock data recovery circuit; PLL - Phase locked loop (clock generation); AFE - Analog front end; TIA - Transimpedance amplifier; DSP - Digital signal processor.

[0045] Figure 1A illustrates an optical switching network according to the disclosure. The optical switching network 1000 includes a plurality of switch devices 100 coupled to a plurality of hosts 200. Each switch device 100 can be configured to establish an optical path from any one of its inputs to any one of its outputs. The switch devices 100 are arranged in a cascaded fashion to form a scalable network. The number of switch devices and hosts may vary. In this example three switch devices 100a, 100b, 100c and three hosts 200a, 200b and 200c are represented. Each host includes at least one network interface controller (NIC). For example, a host may contain one or more processors such as one or more processing units. Various processing units could be considered including one or more of a graphics processing unit (GPU), a central processing unit (CPU), an optical processing unit (OPU), and a Tensor Processing Unit (TPU), to name a few. In addition, the host may include a memory.

[0046] In this specific example the switch device 1, 100a is directly connected to the host 1 200a and host 2 200b, and switch device 3 is directly connected to host 3 200c. It will be appreciated that the connections between individual switch devices and individual hosts may vary. For instance, the switch device 2 may be connected to additional hosts; the switch device 1 may be connected to all three hosts 2001, 200b and 200c, etc.... A given host may also be connected to several switch devices. For example, the host 200a may be connected to all three switch devices 100a, 100b and 100c. A switch device may be connected to multiple other switch devices with single or multiple links. In figure 1A the switch device 100b is connected to both the switch devices 100a and 100c.

[0047] The optical switching network 1000 has a network data plane and a network control plane distributed among the plurality of host and the plurality of switch devices forming the network. Stated another way the network data plane is formed by the host data planes and the switch data planes of all host and switch devices present in the network. Similarly, the network control plane is formed by the host control planes and the switch control planes of all hosts and switch devices present in the network.

[0048] The network control plane is realized through point-to-point connections between two given devices on the network (i.e. host to switch or switch to switch) and thus forms a multi-hop network. A "hop" refers to an optic- electric-optic conversion on a device. These conversions happen on the control plane only. The network data plane remains point-to-point between two distinct hosts only since the switch devices on the network will not inspect the data plane traffic. The data plane is all-optical and just enables two hosts on either end to form a point-to-point direct optical link.

[0049] The network interface controllers are configured to provide network protocol information to the network data plane and control information to the network control plane for routing and flow control of optical data to be transmitted between a source host and a destination host among the plurality of hosts.

[0050] The control information may include: routing information, that is the ID of a source host and a destination host. It may also include various control character such as a request, a cancellation, an acknowledgement (ACK) and a non- acknowledgement (NACK), among other network state related information.

[0051] Therefore, the term control information refers to any type of control characters and control plane state information as well as routing and device state information exchanged on the control plane of the optical switching network 1000. The optical switching network control plane uses serial data transmission (e.g. over an optical sideband wavelength on the same fibre that also carries the wavelengths used for the data plane). The devices on either end of such a control plane communication link negotiate the state of the link (link bring up) to then exchange information about routing requests and acknowledgements or delayed acknowledgments as well as device discovery and enumeration information.

[0052] Essentially, the protocol on the serial control point to point link is a low effort L2 protocol which can carry the control data required to set up an optical path from the source host to the destination host through one or a cascade of switches. An acknowledgement ACK, may be provided if the all- optical data plane connection requested by the source host has successfully been established by all switch devices in between the source and destination hosts. The exact encoding of this information is not specified and subject to performance and implementation considerations. Control characters include ACK, NACK, delayed ACK and requests, to name a few. Additional control information the characters may carry are "destination host” for requests or "time stamp for next available route” in delayed ACKs.

[0053] The control information can also include a target address which can then be translated into a smartOCS meta network physical address by the SCP. Other potential information that could be forwarded are traffic priority classes encoded in the higher-level protocol to interrupt ongoing transmissions. This information is explicitly passed into the SCP to avoid any required inspections of network protocol information on the SDP.

[0054] The network protocol information may vary depending on the specific use case. For instance, when using optical switching network 1000 in conjunction with an Ethernet environment then, the network protocol information would be the L3 data (IP packets) to be sent through the network 1000. Stated another way the network protocol information contains all information as provided by the network protocol layer (OSI layer 3). However, the network protocol information is not to be taken synonymously with the IP packets from L3, because the optical switching network 1000 can choose to "bunch” packets together into a larger transaction which then form a stream of data on the SMN data plane. From this perspective, the smartOCS meta network ignores the packet characteristic of the L3 data. If another network protocol is used (other than Ethernet) such as UALink, PCI-Express or Infiniband, the network protocol information tunnelled through the data plane of the optical switching network 1000 would be the payloads of those protocols which layout is different from the layout of the Ethernet protocol. The optical switching network 1000 may be referred to as an optical circuit switching OCS network, or smartOCS network. The optical switching network 1000 may be implemented as a so-called meta network, that is as a physical network added to an existing network. For this reason, the OCS 1000 may also be referred to as OCS meta network or smartOCS meta network (SMN). Such a meta network is compatible with different kinds of network stacks on layer L2 and above.

[0055] Examples of protocols that shall be supported by the proposed network include Ethernet, Ultra Ethernet, CXI / PCIe, NVLink or UALink, to name a few. The commonality of these protocols lies in the fact that they are based on point-to-point SerDes enabled PHY layers. These PHY layers can be directly used as data plane PHYs on the host systems with relatively small modifications as described further in this application.

[0056] In operation the control flow takes place on the control plane of the network, while the data flow takes place on the data plane. The control plane and its control plane protocol are responsible for establishing routes through the network on the protocol agnostic, all optical data plane.

[0057] The data plane may be built using optical circuit switch technology. Connections are established between source and destination through their respective PHYs as a physical point to point link through the optical switch matrices of the switches that have to be traversed by the light originating from the transmitter of one host to the receiver of the other host. All connections on the data plane are taken to be bidirectional in nature throughout this document. An option is also presented, describing how the data plane connections can be used unidirectionally. Since the optical circuit switch technology is protocol agnostic, the bidirectional connection may be established on several channels (lambdas). It is assumed that the number of wavelengths (lambdas) will be the same in either direction. The control plane can be built based on varying point to point physical interconnect technology. Here, each host of the network is attached to one or more switches of the network through a link. In contrast to the data plane, the control plane will never be interrupted or disconnected. Switch events of the switch will only ever affect the physical connection of a host with another host on the data plane. More specifically, if the control plane connection between a host and a switch device malfunctions or becomes disconnected, the host is considered non-operational.

[0058] The optical switching network of the present disclosure permits fast routing reconfiguration. For instance, the network may reconfigure routing in less than about a microsecond.

[0059] Figure IB shows the network layer definition as put forward by the Open Systems Interconnection (OSI) model. The optical switching network 1000 of figure 1A (the meta network portion) implements the physical and parts of the coding sublayers. This allows it to be used with the network protocols mentioned above which implement the packet and transport layers and any layers required beyond them.

[0060] Figure 2A is a diagram of a switch device for use in the optical switching network of figure 1. The switch device 100, also referred to as smart optical circuit switching (smartOCS) device, includes an optical switch matrix 140 coupled to a plurality of ports I I OI-N. For instance, the number of ports may be an integer N, usually but not limited to a power of two. All ports 110 are implemented in the same way and therefore identical to each other. It will be appreciated that the ports I I OI-N may be arranged in different ways in the switch device. Figure 2A shows a possible connection between a specific port 110k and another specific port l lOj. The optical switch matrix 140 provides an optical input and an optical output for each one of the N ports. Several optical multiplexers MUX 170I N and demultiplexers DEMUX 130 I-N are provided. One MUX and one DEMUX are provided per port 110 of the optical switch 100.

[0061] An electronic control plane 160 and a reference oscillator 180 are also provided. The electronic control plane 160 includes N control path circuits 161 I-N, whose fundamental operation clock is provided by one or more PLLs 162; and a control network layer 163. The control network layer 163 includes an arbitration functionality, routing tables and usage statistics, monitoring and failure recovery functionalities. The electronic plane 160 of the switch device is part of the control plane of the network.

[0062] A port 110 may be implemented as the combination of an optical input 111 with an optical output 112. There are at least two distinct optical channels encoded onto the fibre which connects to a port, one for data and one for control.

[0063] Alternatively, a port 110 may be implemented as the combination of an optical input, an optical output, an electrical input and an electrical output (In this case, the electrical input may be directly connected to RX and TX of port control path 161 for port K). There can be a single or multiple optical data channels encoded onto the fibre / waveguides passing through the optical switch matrix 140.

[0064] The control and data channel(s) enter the switch device 100 through the port 110, the optical input is connected to an optical demultiplexer 130 which separates the control channel from the data channel(s). The data channel(s) is / are routed through the switch matrix 140 (in this example and without loss of generality, to port HOj). The optical control channel signal is converted to the electronic domain via one of the N photodiodes 150I N or via one of the N optical pattern recognition circuits 120I N. The control channel signal is then directed to the electronic control plane 160, more precisely to the control path 161 assigned to the given port. After passing through the switch matrix 140, the optical data are combined with control information from the control path 16 lj at the multiplexer 170 j.

[0065] Figure 2B is a diagram of another switch device for use in the optical switching network of figure 1. The switch device 100’ of figure 2B is similar to the switch device 100 of figure 2A with some modifications, and corresponding components are represented with the same reference numerals.

[0066] In this implementation the ports I I OI N include the optical output 112, the optical input 111, as well as the electrical input 113 and the electrical output 114. For each port 110, the optical input 111 and the optical output 112 are connected to the optical switch matrix 140 via an optical connection. Similarly, the electrical input 113 and the electrical output 114 are connected to the electronic control plane 160 via an electrical connection.

[0067] In both figures 2A or 2B, the electronic control plane 160 could be either an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or a system on chip (SoC). The switch matrix 140 and ports on the data plane are all optical without buffering so that communication through the network happens all-optical or regenerated-optical from the source host to the destination host. The buffering is delegated to the hosts connected to the switch device.

[0068] The procedure of establishing a path through the network on the data plane is handled by the control plane. The control plane ensures appropriate signal flow control through the data plane. All connections on the control plane within the network are point-to-point communication paths. Possible connection scenarios are host-to-switch or switch-to-switch. The electronic control plane 160 may be adapted to generate a clock reference based on a reference oscillator 180. The clock distribution to other devices in the network happens by embedding this clock in the control plane point-to-point link protocol. An appropriate number of transitions will be guaranteed by choice of datalink layer protocol encoding on the control path.

[0069] A switch device 100, 100’ can be either a clock generator or a clock follower. Every network has a single clock generator. The clock generator is determined during network initialization. All other devices are clock followers. The hosts 200 connected to the network are always clock followers. If a switch device 100 is a clock generator, it uses its high precision, internal quartz-based reference oscillator to drive the phase locked loops (PLLs) 162 of all its control plane PHYs 161 I N. In turn these PHYs are all connected to the control plane PHYs of the attached host systems through the control channel. The switch reference oscillator 180 becomes the timing reference for the attached hosts as follows. The switch PLLs 162 are all locked to the local reference oscillator 180. The PLL clock outputs are used to drive the multiplexer circuitry of all PHY transmitters of the switch.

[0070] The PHY transmitters are responsible for serially communicating the control plane protocol contents originating from the switch control plane to the hosts attached to a given switch port. In this way, the data stream sent on the control plane from the switch to each host also contains phase and frequency information that is locked onto the switch devices’ own reference oscillator. How this information is used to form a source synchronous link on the data plane is described in the sections which describe the host device architecture.

[0071] Figure 3A is a partial diagram of a host NIC for use in the optical switching network of figure 1. Figure 3B illustrates an exemplary implementation of a transceiver for use in the circuit of figure 3A. Figure 3C shows an example implementation of a PLLs and clock distribution circuit for use in figure 3A. Among other entities, the NIC contains a MAC, the network PCS and PHY (205) and a PMD. The PHY 205 is made of a NIC data plane 210 and a NIC control plane 220, also referred to as smart NIC data plane SDP and as smart NIC control plane SCP. The NIC data and control planes 210 and 220 are coupled to an optical physical medium dependent PMD layer 230. The PMD layer 230 is configured to convert a signal from electrical to optical on TX and vice versa for RX. On the data plane, this is always true. On the control plane, conversion between optical and electrical domain may also be omitted, if the host is electrically connected to a switch device on the control plane.

[0072] The NIC data plane 210 contains a physical coding sublayer PCS 211, coupled to a data physical sublayer, data PHY 212. The NIC control plane 220 contains a set of buffers 221, a control PCS 222, a control PHY 223, an address translation table 224, and a PLLs and clock distribution circuit 225. The address translation table 224 is used to translate between the higher level L3 protocol address and the SMN (L2) address.

[0073] Therefore, the NIC control plane 220 has both the control plane PHY 223 and the control plane PCS 222. The control plane PCS 222 is responsible to provide the control information to 223 with the appropriate line coding and packaged into information characters such that this information can be readily decoded by the PCS layer within the switch device attached to the other end of the control path (i.e. PCS of 161 in Figure 2A). The control plane PCS 222 is coupled with the data plane PCS 211 and the MAC layer to process network routing information and carry it through the control plane network. This allows the optical switching network to never have to evaluate information contained in the packet data stream on the data plane. Thus, irrespective of the transformations within the data plane PCS 211 demanded by a given standard (such as scrambling or recoding), the control plane always has unhindered access to all routing and packet information required to perform its routing tasks.

[0074] The hosts 200, which are part of the optical switching network 1000 (also labelled smartOCS network) use a NIC dedicated to both the switch devices 100 (also labeled smartOCS devices), and the desired network protocol such as Ethernet, CXL, among others. This is achieved in the implementation of the NIC software layers (beyond the hardware) and the MAC layer (L3) which is part of the NIC hardware. In this document, only the technical details of the network protocol agnostic functionality of the switch devices are discussed. Wherever required for clarity, a potential Ethernet implementation is described. It will be appreciated that Ethernet is just one example of a network protocol to be supported by the switch devices and the optical switching network they form.

[0075] The NIC (smartNIC) of a host 200 is configured to provide the network protocol information to the data plane of the optical switching network 1000 and the control information for routing and flow control of the data to be transmitted to the control plane of the optical switching network 1000. The NIC could be physically located in a dedicated plugin card (Peripheral Component Interconnect Express PCI-E CEM) or be fully integrated into a SoC.

[0076] Irrespective of its location, the smartNIC implements the full stack required by the given network protocol down to the layer of packetization. In Ethernet, this would be the medium access (MAC) layer. The smartNIC may also be compliant with layers below the protocol layer, such as the coding layer (again, in Ethernet terms: the PCS). However, if a network protocol defines a clear-cut interface between layers of packetization and coding, the smartOCS NIC may opt to implement a coding sublayer deemed optimal for the optical engine it needs to support. In Ethernet language, the optical engine would be represented by the physical medium dependent (PMD) layer.

[0077] Data entering the host NIC and arriving at the interface from packet layer to coding layer will still be in its native binary form, unscrambled and unencoded. Instead of the plain PCS and physical sublayer (in Ethernet terms called physical medium attach or physical medium attachment - PMA), the smartOCS NIC possesses two PCS layers 211, 222 and two PHY layers 212 and 223. One for the data plane and one for the control plane.

[0078] The host NIC control plane 220 is configured to perform one or more of the following tasks: i) establish flow control, ii) perform arbitration through the optical switching network, iii) perform network address translation if not done through the packet layer, iv) perform fast relock operation upon smartOCS switch events.

[0079] The data physical sublayer (212) includes a clock data recovery circuit, and the NIC control plane (220) is configured to store phase accumulator states of the clock data recovery circuit of the data physical sublayer (212) for a plurality of hosts. The control PCS circuit 222 may perform flow control by creating backpressure upstream towards the packet layer by not accepting data at the smart NIC control plane SCP interface and not providing valid data towards the MAC.

[0080] Arbitration through the optical switch network 1000 is performed within the switch devices 100 by the arbitration module of the control network layer 163. Specific control characters are sent to the switch device(s) 100 attached to the host 200. These characters inform the switch device of a particular connection on the data plane that the host wants to establish (a request). In turn, switch 100 informs a host about the availability or unavailability of a requested connection through switch matrix 140 by sending acknowledge characters or non-acknowledge characters on the control network layer. If a path through 140 is requested by one of the hosts 200 connected to a switch device 100 which is already in use, the process known as "arbitration” is performed inside of the control network layer 163 to decide, in what order and at what time pending request are to be served (i.e. acknowledged).

[0081] The protocol circuit 224 is configured to perform network address translation. The destination ID information, which may also be referred to as label, can be provided by the MAC layer or be extracted from the packet header before scrambling / encoding in the data plane PCS 211 if the network protocol layout is known and documented. When the destination ID information is provided by the MAC layer, a dedicated interface provides a target host ID in a format common to the network protocol. This host ID is then translated to a network specific label which is then used to route through the control network layer 163 of the switch device 100, to set up the path through the data plane.

[0082] The label may be a global label or a sequence of local labels. A global label is valid throughout the network and used to find the correct physical path through the network in each and every switch device 100 of the network using switch device local routing tables. This approach scales well but requires appropriate network initialization through a network management instance. In the sequence of local labels, every label is valid for finding the path through the next switch device. This is a global routing scheme, and all hosts 200 need a view of the network topology and associated host IDs. This is mostly suitable for small networks and keeps the complexity of routing tables very low within the switches.

[0083] In order to perform fast relock on the SCP, the NIC control plane 220 and the control plane 160 of each switch devices in the network leverage fast SerDes relock times upon reconfiguration of optical links on the data plane through multiple different features on transmitting and receiving side. The control path between a switch device 100 and a host 200 connected to it is always active and constantly exchanging control and / or idle characters. This allows the transceivers on either side of the control link to maintain lock at all times and have a recoverable reference clock at their disposal. The clock data recovery (CDR) of the receiver in the NIC control plane 220 locks onto the clock embedded in the data stream that originates from the switch device 100 it is connected to. The CDR then outputs both the recovered data (going towards 222) and a recovered clock. The recovered clock is used as a reference to PLLs and clock distribution circuit (225) which performs some jitter cleaning on the clock and provides reference clocks to all circuits present in the Data PHY 212.

[0084] This switch device may either be a clock follower or a clock generator. In any case, the clock embedded in the control channel data stream will be frequency locked to the network base reference oscillator.

[0085] As shown in figure 3C the PLLs and clock distribution circuit 225 may be implemented with two PLLs. The first PLL (PLL_1) locks to a local reference oscillator on the host. The second PLL (PLL_2) locks to the CDR recovered clock of the control channel, which then drives the data plane transmitter and receiver. This arrangement makes the data plane transceiver source synchronous.

[0086] Clock coherency refers to a situation in which a transmitter and a receiver which form the two ends of an optical link operate based on a same timing reference (reference oscillator). This means that transmitter and receiver will never experience a difference in momentary frequency (frequency as measured over a finite time interval) but will only differ in relative phase. Extrapolated to an entire data plane, this means that _any combination of transmitter / receiver pair is clock coherent.

[0087] The transmitter, receiver and CDR circuit of the NIC control plane 220 (see CDR in Control PHY 223) is clocked by (or driven by) the first PLL (PLL_1). As such, the control channel is not source synchronous and requires both frequency and phase tracking in the CDR. The recovered clock of the CDR (of 223) is passed on to a jitter cleaning PLL (second PLL, PLL_2) which drives the transmitter, receiver and CDR circuit of the NIC data plane SDP 210 (See arrow between circuit 225 and CDR of Data PHY 212 in figure 3). Since the same is true for any other host 200 in the optical switching network 1000, any established data channel through the network between two hosts will be a source synchronous serial link and will not require frequency acquisition. As a result, the phase and frequency accumulators in the CDR circuits on either end of the data channel will not have to recover frequency when their receivers lock onto the data stream of the opposing sides transmitter after a switch operation thereby decreasing overall relock time.

[0088] Equalization caching is performed on the data plane of the network. At least one host 200 in the network (or possibly all the hosts) may be provided with an equalization circuit adapted to cache equalizer settings for a pair of transmitter / receiver on each side of a physical link associated with a specific connection in the network, and to recall such settings when the same connection is established.

[0089] The Data PHY 212 of the host 200 has a transceiver circuit. When two transceiver circuits form a physical link, their equalization circuits try to remove as many physical signal distortions as possible. The equalizer setting that a transmitter (tx) / receiver(rx) pair on each side of a physical link will train to are specific to that particular pair of tx / rx. This equalizer training process occurs inside of the tx / rx circuitry and may be relatively timeconsuming.

[0090] The host 200 is provided with a register that stores equalization states or settings. For instance, the register may be located in the PMD 230 or the data PHY 212. The host caches the last known equalizer settings for a pair of tx / rx, (that is already trained or "good” settings) and can recall them the next time the same connection is established. As a result, equalizer training can be skipped thus saving reconnection time.

[0091] The NIC control plane 220 receives regular notifications from the switch device 100 it is connected to. These include information on when a reconfiguration will take place and to which other host on the network the NIC control planes will be reconnected to. By storing the last known phase accumulator states for all hosts of the system a particular host has been connected to in the past, this phase information can be restored prior to the actual reconfiguration taking place and the CDR of the SCP can be frozen (disabled phase updates). Upon physical reconnection due to a switch event in the switch devices 100, the number of required phase updates to arrive at the optimal sampling location will at most be a few steps (the number of required steps usually is not zero due to inevitable temperature and voltage drifts in all devices of the network). In any case, time is saved for reestablishing a connection.

[0092] The same procedure described above for caching and recalling the phase accumulator state of the CDR can also be applied for the equalization state in transmitter FIR and receiver FFE / DSP for optimal bit error rate. In multichannel / lane data plane systems (wave division multiplexing, polarization multiplexing and mode multiplexing), in addition the host specific intra-channel / lane skew can be cached and recalled.

[0093] The optical switching network may also be configured to perform loopback on disconnect. Time constants in a transceiver based serial link are usually on the order of THz (optical bandwidth), GHz (analog bandwidth) or hundreds of MHz (CDR, DSP, line coding). Yet, disconnecting an optical link and reconnecting it is reported to take a long time, sometimes in the order of milliseconds (kHz). A likely candidate for long down times if PLLs are not powered down are the baseline wander effects due to inactive; AC coupled links and the slow current sources in analog frontends that are required to calibrate a receiver to the optimal DC biasing points. To avoid this behavior if a host needs to be disconnected from the network data plane (because at this point in time the host has no communication partner on the data plane), the smartOCS will direct a host without current communication partner to be put into a loopback configuration through the switch devices optical switch matrix. In this context "disconnected” simply means that a particular host on the network does not have a communication partner on the data_ plane, however its point-to-point connection to the switch device always stays active and alive.

[0094] The optical output of the host is redirected through the switch matrix to its own optical input thus forming a loopback on the data plane. In this way, the biasing sensitive circuitries inside the Optical PDM 230 will always see an active link, even if only idle characters are transmitted across that (loopback) link. Note that the point-to-point connection between the host control PHY and the electronic control plane of the switch device it is connected to always stays alive.

[0095] Turning to data flow control, when data is available for transmission at the MAC layer, the optical communication switching network 1000 considers this interface the last interface at which on-chip flow control is realized. The serial links on the data and control planes transmit regularly changing symbols at every unit (limited run length). This is to maintain the CD Rs of the receivers in lock and limit the magnitude of baseline wander.

[0096] To this end, in addition to scrambling the data to be transmitted (by manipulating the data stream before transmitting), the coding layers of data and control planes will introduce idle characters into their data streams whenever there is no valid data available from the packet layer (SDP) or no control information needs to be transmitted (SCP).

[0097] On the receiving side of the data plane, these idle characters are used by the coding sublayer to monitor the health state (absence of errors) of the link but are not passed on to the packet layer. They are dropped from the data stream and thus, there is a clear mechanism to indicate the availability of useful data at the interface between coding to packet layer.

[0098] The NIC control plane 220 can interfere and override the way in which the coding sublayer reports data availability to or readiness to accept data from the packet layer. In this way, the NIC control planes 220 in the two hosts forming a direct link across the data plane establish the flow control with the packet layer. Additionally, the NIC control plane 220 conceals the actual physical link availability during switching operations. Usually, if an optical circuit switch breaks the connection between two hosts A and B and establishes a new connection from host A to C, the physical layer in host A would, for a certain period, indicate a loss of link to the coding sublayer and the coding sublayer in turn would then report a link unavailability to the packet layer. This unavailability indication would then propagate all the way to the operating system which would consequently free all packet buffers in memory associated with the network link. Upon successfully establishing the physical connection between hosts A and C, the physical layer would indicate the presence of a valid signal to the coding sublayer which would then again try to reestablish lane deskew, descrambling lock and FEC lock. Once this has been accomplished, the coding sublayer would then report link availability to packet layers and above. As a result, the operating system would reinitialize all transmission and reception buffers in memory and make the network link available to applications again. This process can easily take several milliseconds and would render all link re-establishment speedups implemented by the NIC control plane 220 and the NIC data plane 210 superfluous.

[0099] The optical switching network 1000 may be configured to avoid operation system (OS) interaction altogether and perform a switching operation on layers at and below the physical coding layer only. This can be achieved through network protocol dependent tweaks to all network layers above the coding layer combined with the ability of the NIC control plane to hide the fact that a physical link is actually "lost” during a smartOCS switching operation. This may be accomplished by not passing link status updates to the upper layers. Only in cases when the NIC control plane or the NCI data plane run into irrecoverable hardware errors will layers at and above the packet layer be involved and made aware of a link outage.

[0100] To improve performance, the NIC control plane 220 may be provided with a set of buffers (221), also referred to as per-destination-host packet buffers to collate a set of smaller transactions if necessary. This is because the network efficiency depends on the duration of data transmission to and from a given host. If data sizes and thus duration of transmission is large, the time to reconfigure to connect from one host to another decreases in importance. The network efficiency Neff is calculated as:

[0101] Neff = t_dtran / (t_dtran + t_overhead) in which t_dtran is the data transmission time interval, and t_overhead is the overhead time that includes the time elapsed for all tasks required to form a new connection between two hosts on the data plane.

[0102] To keep the number of host packet buffers manageable, a least recently used (LRU) algorithm may be used to reassign packet buffer space to new host IDs. Since the optical switching communication network 1000 is primarily targeted at applications with regular traffic patterns, it is expected that an upper bound for the number of required buffers can be determined for a given cluster and workload size.

[0103] Data presented to the NIC data plane 210 at the MAC / PCS interface is indicated via a valid signal. For protocol stacks where this is not customary (such as Ethernet, although more recent standards now employ the reconciliation layer (RS) between MAC and PCS), the MAC layer needs to be adapted accordingly. Standards like Ethernet define a physical interface standard such as 10 gigabit media-independent interface (XGMII) to connect lower-level network layers. An Ethernet example would be the connection of a MAC with PCS in one physical chip connecting to another PCS / PMA / PMD in another chip such as a transceiver chip on an XSFP module.

[0104] Interfaces like these would need to be amended to be compatible with the switch devices 100. This is not only because of the requirement of flow control at the MAC / PCS interface but also because the MAC layer may need to present the network host ID to the NIC control protocol directly. This is because the propagation of routing information may happen according to multiple different scenarios.

[0105] For Ethernet like protocol - the physical interface presents data that may already be scrambled to the NIC data protocol. In this case, the header and thus the routing information of a packet cannot be extracted without descrambling. This would make the NCI data protocol very network protocol dependent and optical switching communication network 1000 seeks to keep the NIC control and data planes as network protocol agnostic as possible.

[0106] For modified Ethernet like protocol, several cases may be considered. In a first case, the MAC does not perform IP network address to network destination ID translation. The routing tables in IP based networks are kept in the network switches and are updated through various protocols (e.g. ARP). The optical switching communication network 1000 cannot rely on mechanisms like these as the underlying data plane is a circuit switched network which only establishes point to point connections between hosts. As such, the destination ID must be known at the initiating host side already.

[0107] In a second case, the MAC is modified to perform IP address to network destination ID translation. It must then present this destination ID to the NIC control protocol, thus altering the way in which the MAC / PCS interface in network stacks like these are defined (and also how XGMII based interfaces work). In a third case, memory semantic fabrics (like CXL / PCI-E or NVLink) - here, the base memory address of an address range is used either directly or hashed as the destination ID. In resemblance to the second case, this translation may either happen on the protocol layer or within the NIC control protocol.

[0108] Figure 4 is a diagram of an optical switching network comprising several bridge devices. The optical switching network 400 may be referred to as an optical circuit switching OCS network, or smartOCS network. The optical switching network 400 includes a plurality of switch devices, also referred smartOCS switch devices as described above with reference to figures 2A and 2B. The optical switching network 400 has several domains associated with one or more smartOCS switches. For instance, the domain 410 is associated with the smartOCS switch 411, and the domain 420 is associated with the smartOCS switch 421 etc.

[0109] The smartOCS switches of one domain are coupled to the smartOCS switches of another domain via a bridge device, also referred to as smartOCS bridge device. For instance, the domain 410 is coupled to the domain 420 via the bridge device 401, and to the domain 430 via the bridge device 402, respectively.

[0110] A domain may have one or more switches and no or a plurality of hosts, and each host may have a specific ID, also referred to as label or address, in that domain. For instance, a destination host, also referred to as target host, may be known as ID = 3 in the domain 420 but as ID = 5 in domain 430. The bridge device may be provided with an address translation circuit to translate the label. This may be achieved using a programmable look-up table. The look up tables are thus a part of the SMN control plane. The bridge devices may be used for optical signal regeneration, buffering, link monitoring and head of line blocking mitigation.

[0111] Figure 5 is a diagram of a bridge device for use in the optical switching network of figure 4. The bridge device 500 includes a first port circuit 510, a second port circuit 520, a buffer circuit 530, a state machine 540, and a memory 550 for storing data during buffering. The memory 550 may be shared among multiple bridge devices 500 on a single physical die, but it is specific to the bridge device and not shared with the network. A fall through path, also referred to as bypass path, may also be provided, either as part of the buffer circuit 530 or as a separate path.

[0112] The state machine 540 is provided on the control plane and is configured to receive routing information for routing a data stream between a source host and a destination host in the optical switching network and to instruct the buffer circuit 530 to either store the received data stream or to pass the data stream to the second port circuit 520. This is based at least on the routing information, and on buffer state information indicating whether the buffer is full or has available memory space. This may also be based on transmitter information indicating if the data plane transmitter 522a is available for transmitting data. The state machine 540 may also instruct the buffer based on control information exchanged on the control plane of a smartOCS meta network.

[0113] Optionally, the bridge device 500 may also include a translation circuit 542 configured to perform network address translation. For instance, the id / address of a host in one domain may be translated to a different id / address in another domain. This provides flexibility in setting up multiple independent address domains within a single physical network (SMN).

[0114] The first port circuit 510, also referred to as input port, is formed of a receiver physical medium attach PM A RX 511 coupled to a data plane receiver 512a and control plane receiver 512b. The second port circuit 520, also referred to as output port, is formed of a transmitter physical medium attach PMA TX 521 coupled to a data plane transmitter 522a and a control plane transmitter 522b.

[0115] The PMAs are adapted to perform the opto-electric and electro-optic conversions and to provide a mechanism to attach the fibre(s) to the input or output port. The PMAs 511 and 521 may be implemented in different ways. For instance, a PMA may include photodetectors, optical demultiplexers and optionally semiconductor optical amplifiers SOAs and Mach-Zehnder-Modulators MZM (or another form of optical switch) for optical path steering and the TIA for both data plane and control paths. The receiver circuit 510 and transmitter circuit 520 on physical and link level operate like their counterparts in host endpoint NICs. The difference on the physical layer lies in the way in which an optional SOA(s) may be used for preamplification at the receiver input or bypassed by means of a set of MZM devices or other form of optical switch.

[0116] Additional optical PMDs (not shown) are provided in the bridge device 500 to perform the optical to electronic data conversion for data received at 512a / b, and the electronic to optical conversion for the data arising from 522a / b. The opto-electro-optic (0E0) conversions may also be used to achieve: i) optical signal cleaning in multi-hop scenarios; and ii) head-of-line blocking reduction and more efficient routing path reservation through electronic buffering.

[0117] The buffer circuit 530 includes a per-source / destination buffer. A per source / destination buffer has a specific memory area allocated to hold data transferred between a particular sender (source) and receiver (destination). The data stream buffering capability may be used to avoid head-of-line blocking in all optical network data planes of smartOCS meta networks. The bridge device can only transmit data to a single target host ID at a time. This is due to the circuit switched nature of the smartOCS meta network. If there is a lot of data stored in the memory 550 for target host ID 3, the bridge device gets the grant from the switch(es) connected to its TX side 520 and can start transmitting data to host ID 3. In the meantime, additional data may arrive at receive side 510. Let’s suppose one get data from source host 0 for some time with target host 4 followed by a pause followed by more data from hostO for some time for target host 4 again. During all this time, the bridge device is transmitting to target host id 3. However, since the bridge device has per-source / destination buffering, all the data coming into the bridge device for target hosts 4, can be buffered up. As a result, the once isolated transactions from source hosts 0 and 1 would be collated by the bridge device in its buffer to then be forwarded as one block to target host 4 once that route has been granted on the bridge device TX side.

[0118] In operation, once a link has been established by the control plane, data from the fibre enters the bridge device 500 through the physical medium attach PMA 511. The control data is passed into the control plane receiver port P0 SCP RX 512b to deserialize the data on the control plane, recover the embedded clock (using a CDR circuit) and process the control characters in accordance with the smartOCS meta network definitions. Control characters may include acknowledgement (ACK), non-acknowledgement (NACK), routing information. Bit deserialization refers to the process of arranging data from a stream of single bits into a block of bits that are processed at the same time but at lower speed (clock cycle). The inverse procedure is called bit serialization.

[0119] A PLL 513 is provided as a clock reference for both data and control plane receiver circuits. The clock recovered by the CDR in the SCP RX 512b is provided to the receiving circuitry in 512a so that the CDR in 512a only needs to do phase adjustment and no frequency tracking of the clock embedded in the data stream of the data path (source synchronous data plane). The relocking support indicated in Figure 5 between SCP RX 512b and SDP RX 512a works like in the host NIC instances. Through phase and equalization caching, the time in which the SDP receiver can relock onto an incoming data stream is reduced.

[0120] The data on the data plane is deserialized per channel in the smartOCS data plane receiver 512a. The data then passe through the buffer circuit 530, or in some cases bypass the buffer to go directly to the transmitter (fallthrough mode). The buffer circuit 530 can either store the deserialized data or operate in a so called fallthrough mode, depending on the information and state decoded and held in the bridge device control plane 501, specifically the state machine 540. The fallthrough mode is active when the incoming data on the SDP RX of the bridge device is associated with a target host ID for which an established route is available on the SDP TX side of the bridge device.

[0121] The state machine 540 is responsible for managing the operation of the bridge device 500. Based on the control information provided by the SCP RX 512b, the state machine 540 instructs the data buffer 530 to either store a data stream received at the SDP RX 512a in the memory 550, or to pass it through to the SDP TX 522a. It does so by keeping record (a copy of the respective register states associated with SCPs) of whether a connection was established on the bridge TX side and to which target host ID the connection is established there. If a connection on the TX side is unavailable, it instructs the data originating from the RX side to be buffered. If there is no buffer space left, the state machine 540 instructs the SCP RX 512b to not accept incoming connection requests in the first place thus creating backpressure on the initiator of the transaction request (another bridge device or a host NIC within the smartOCS meta network). If there are no pending requests for data transmission on the RX side of the bridge device (e.g. no active requests on the RX SCP and no running transactions on RX SDP), the state machine 540 keeps a record of the buffer fill status and requests a route on the data plane to the TX side of the bridge for any of the available buffers. The state machine 540 does this in a round robin fashion and if the switch on the TX side does not have a route available for a given target host ID, the state machine will move on to the next per-src / dest buffer and will ask the TX SCP to request a route for this alternative target ID. Only if there is no data on the RX side of the bridge and all buffers are empty, the state machine 540 will go into an idle state.

[0122] If data is selected for transmission by the control plane state machine 540, then the control plane transmitter port (Pl SCP TX) 522b becomes responsible to negotiate the arbitration across the attached smartOCS switch stages to the next smartOCS receiver. The control plane transmitter port (Pl SCP TX) 522b also establishes flow control for the data plane transmitter port 522a to enable fast optical path establishment and correct protocol padding. Both data plane and control plane ports (SDP TX) 522a and (SCP TX) 522b are adapted to perform data serialization. For instance, 522a and 522b may include a serialization circuit. The resulting bitstream is transmitted to the PMA TX 521 for electro-optic conversion, muxing and injection into the output fibre.

[0123] The amplitude of the data signal transmitted between switch devices decreases at each hop. For large optical switching networks this may result in a significant loss of signal-to-noise ratio (SNR). The bridge devices may be used to "clean” the signal between switch devices, hence improving the reliability of the optical switching network. The opto-electro-optic (0E0) conversion "cleans" the signal by restoring its original digital meaning in the process. Optionally, further amplification of the signal may also be implemented.

[0124] Bridge devices may be coupled together or arranged in different ways to share specific functionalities. It will be appreciated that depending on the application, the bridge device of figure 5 could be implemented without the buffer and the state machine. In this case the ports 510 and 520 would be coupled by a simple link on either end to different network devices.

[0125] Figure 6 is a diagram of a bridge device system for use in the system of figure 4. The bridge device system 600 includes a plurality N of bridge device pairs labelled 611-61N. Each bridge device pair includes two bridge devices that share some features. For example, the bridge device pair 611 includes two bridge devices 611a and 611b in which the state machines 641a and 641b are connected to each other. The bridge device system 600 may be used to improve memory utilization and reduce power overhead.

[0126] In addition, the bridge device 641a may be paired with the bridge device 641b to create a bidirectional bridge device. In this case, the two logical bridge device control planes are coupled through digital logic such that they can implement an optical loopback mechanism through an attached smartOCS switch so that all analog circuitries in the physical layers maintain a state of operation that enables a fast optical relock procedure in the event of a switch reconfiguration event.

[0127] Referring back to figure 4, the only protocol aware part of the bridge device is the control plane circuitry. The smartOCS meta network has constantly running control plane connections which are point-to-point connections between all physically attached devices. If the connection on a data plane between two endpoints of a network through a (series of) smartOCS switch(es) is deactivated, the control plane connections are still maintained. Therefore, the bridge device control plane will always be able to send and receive control information to and from the attached switch device(s).

[0128] A pair of bridge devices sharing information across their control plane can therefore dynamically fill and maintain an address / label translation table (542 in figure 5) for the smartOCS meta network (SMN) control network. As an example, if the control plane implements a label-based routing approach with labels i.e destination ID information) that are valid across a single smartOCS device domain (420 in Figure 4) (can either be a single switch device or multiple switch devices directly connected), the label is only valid within that specific domain.

[0129] Network enumeration refers to the process by which a unique identification (usually an integer value) is assigned to a host connected to a network. In a network enumeration scheme which assigns smartOCS control network labels based on proximity (number of network hops between nodes), the smartOCS switch domains can be set up such that the routing tables within the switch devices are of contiguous nature in the sense that labels within a domain carry meaning as to where (to which link) the data has to be forwarded. In such situations, the address / label translation tables are of compact form and can easily be stored inside of the SRAM of a bridge device to ease the network endpoint enumeration scheme.

[0130] SmartOCS operates on OSI layer 2 which means that the unique identifier is only valid within a single smartOCS device domain (unless bridge devices are used to translate between domains). Unique identifiers on level 1 are sometimes also called "physical address”. Nevertheless, these identifiers are only used at the low OSI layer and have no meaning on higher protocol layers. Higher layers usually use their own addressing scheme and implement a translation table to retrieve the physical unique identification of a target host when passing data from a higher OSI layer to layer 2

[0131] Figure 7 is an example implementation of a bridge device pair. In this example two unidirectional bridge devices are combined to form a bidirectional bridge 700 also referred to as smartOCS meta network bridge.

[0132] The optical PMAs 711 and 721 may be implemented by combining PMARX and PMATX. The smartNIC data and control planes 712 and 722 can be formed by the PO SDP RX 512a, PO SCP RX 512b and Pl SDP TX 522a, Pl SCP TX 522b as shown in figure 5, but by rearranging the components into a form as illustrated in figure 8.

[0133] Figure 8 shows an example implementation of a smartNIC data and control planes (711 or 722 of figure 7). The circuit is similar to the circuit of figure 3A for the host. For ease of reference the same reference numerals are used as in figure 3A, and operation of the circuit can be directly derived from the description of figure 3A. The main difference is that the bridge device does not have a MAC layer, so the control PCS 222 exchanges data flow control information with the state machine 740. In addition, the PLLs and clock distribution circuit 225 receives a reference clock signal from a bridge local oscillator. The local oscillator provided on the bridge may be a crystal quartz oscillator.

[0134] In contrast to a smartNIC in a host, the smartNIC in bridge devices are formed by two smartOCS meta network ports up to OSI layer 2 without any layer 3 (MAC) components and FEC as an optional choice for layer 2. In contrast to the smartNIC in a host, the bridge devices only responsibility is to enable buffering or forwarding of data on the data plane based on the information sent and received and the links negotiated on the smartOCS meta network control plane for its two ports. Given that links are usually bidirectional on the data plane of the smartOCS meta network, a single SMN address translation table for both directions is sufficient.

[0135] A bridge device pair (or bi-directional bridge) as shown in figure 6 or 7 may be configured such that first bridge device performs a primary conversion from a first optical standard to a second optical standard and the second bridge device performs conversion from the second optical standard to the first optical standard. The first standard may encode data using a first predetermined number of wavelengths and / or other optical signalling properties, and the second standard may encode data using a second predetermined number of wavelengths and / or optical signalling properties which may differ from the first. For example, the first standard may be 400G- FR4 and the second standard 400G-DR4 or the first standard may be IEEE standard and the second standard a custom optical transmission specification.

[0136] Figure 9 is a diagram of an exemplary physical architecture of a bridge device. SmartOCS bridge devices are primarily defined by their overall logic structure. Bridge devices can be implemented physically in different ways. The bridge device 900 is a specific example of a bridge device implemented as co-packaged electronic integrated circuit (EIC) 910 and photonic integrated circuit (PIC) 920 solution. Both EIC and PIC are packaged back- to-back using flip chip technology through a substrate (either silicon or organic for example) 930. The use of through silicon vias (TSV) or direct vias leads to low connection inductances and therefore enables high bandwidth signal transmission from PIC to EIC and vice versa.

[0137] The substrate 930 is connected to the PIC 920 and EIC 910 via a plurality of solder balls 940. The common substrate 930 also permits to integrate an additional memory chip such as high bandwidth memory (HBM) to allow for large transactions to be buffered by the device for many outstanding transactions targeting multiple network endpoints at the same time. This second level memory may be paired with a first level SRAM cache and a cache eviction strategy into the HBM. A pair of fiber attached units (FAU) 950 is provided on the PIC 920 to connect the bridge device to a set of fibers. It will be appreciated that the bridge devices may be implemented using other packaging schemes.

[0138] Figure 10 is a diagram illustrating communication between two host devices using two switch devices connected to each other. When two hosts A and B try to establish an optical path on the smartOCS data plane, this request is communicated through the series of switch devices provided between the hosts. In this example the network does not include any bridge devices.

[0139] In operation a contiguous node id enumeration scheme is used and the routing tables within the smartOCS switch control planes define the correct port for a given network topology. This is achieved by i) assigning ranges of valid node ids to a particular port in cases when a switch device is connected to another switch device or ii) assigning a single node id to a switch port in the routing table if only a host endpoint is connected to the given physical port. When host A with node id 0 requests a path to host B (knowing that it has node id 32 for example), it does so by creating a routing request on the control plane with the added information that it wants to communicate with node 32. Without loss of generality, the example provided here assumes smartOCS switch devices with 32 ports.

[0140] The request is sent to the control plane of switch device 0 and a lookup is performed on the routing table. Switch device 0 determines that in order to fulfill the request, an optical path has to be established between port K and port L within its optical switch matrix and that the request has to be forwarded on the control plane to port L. This is because the routing table for port L has a node range associated with it which indicates that on port L is attached either another switch device or a bridge device instead of a host endpoint.

[0141] Switch device N then receives the routing request on its control plane on port R and will determine, through routing table lookup, that port V is connected to the desired node 32 for the present routing request. If the path is available, it will set up the route on the all optical data plane and inform switch device 0 that it too can establish the optical path between ports K and L now to form the all-optical connection between host A (0) and host B(32) on the data plane (we neglect the additional communication to support fast optical transceiver relock in the description here and defer to the smartOCS meta network patent description instead). As can be seen, with a growing smartOCS device domain, the node id fields would start to grow larger, and the routing lookup tables too would grow in complexity.

[0142] Figure 11 is a diagram illustrating communication between two host devices using two switch devices connected to each other via a bridge device pair. Figure 11 illustrates the same situation as described above. However, instead of a single smartOCS device domain, the bridge device separates the network into two distinct domains: a first domain associated with the host A and switch device 0, and a second domain associated with host B and the switch device N.

[0143] The node ids only carry meaning within a given domain and are translated by the bridge device lookup table. Most if not all network topologies have well defined symmetries and it is along these symmetries that this kind of domain and label separation works very well.

[0144] In this specific example the host A with node id 0 requests a path to host B, also with node id 0. The request is sent to the control plane of switch device 0 and a lookup is performed on the routing table. Switch device 0 determines that in order to fulfill the request, an optical path has to be established between port K and port L within its optical switch matrix and that the request has to be forwarded on the control plane to port L. The bridge device performs an address translation such that the id 32 in the first domain turns into id 0 in the second domain. It will be appreciated that the address translation scheme presented is just one example of how the additional flexibility of a bridge device can be exploited. Switch device N then receives the routing request on its control plane on port R and will determine, through routing table lookup, that port V is connected to the desired node 0 for the present routing request. The bridge device with its separated control and data planes still maintains the protocol agnostic aspect of the smartOCS meta network while providing additional flexibility at very low latency and the added benefit of optical signal SNR cleanup on the data plane.

[0145] A bridge device might choose to implement a FEC or a set of different FEC mechanisms to support various standards if beneficial from a system design point of view. However, it is the only mechanism above OSI layer 1 which would be supported by the device.

[0146] Although the bridge device has been described for use with an optical switching network, it will be appreciated that the bridge device can be used within an optical link provided between two host devices. In this case the bridge device may be used to perform conversion between different optical standards. It may also be used to extend the physical range of the optical link.

[0147] A skilled person will therefore appreciate that variations of the disclosed arrangements are possible without departing from the disclosure. Accordingly, the above description of the specific embodiments is made by way of example only and not for the purposes of limitation. It will be clear to the skilled person that minor modifications may be made without significant changes to the operation described.

Claims

CLAIMS1. A bridge device for use with an optical switching network or an optical link, the bridge device comprising a first port circuit for receiving a data stream; a second port circuit for transmitting the data stream; a buffer circuit and / or a bypass path provided on a data plane.

2. The bridge device as claimed in claim 1, comprising a state machine provided on a control plane; the state machine being configured to receive routing information for routing the data stream between a source host and a destination host in the optical switching network or optical link, and to instruct the buffer circuit to either store the received data stream or to pass the data stream to the second port circuit, based on the routing information, and buffer state information indicating whether the buffer is full or has available memory space.

3. The bridge device as claimed in claim 1 or 2, wherein the buffer circuit comprises a per source / destination pair buffer.

4. The bridge device as claimed in any of the preceding claims, wherein the first port circuit comprises a first physical medium attach coupled to both a data plane receiver and a control plane receiver; and wherein the second port circuit comprises a second physical medium attach coupled to both a data plane transmitter and a control plane transmitter.

5. The bridge device as claimed in claim 4, wherein the state machine is configured to receive transmitter information indicating if the data plane transmitter is available for transmitting data.

6. The bridge device as claimed in claim 4 or 5, wherein the control plane receiver is configured to receive instructions to perform one or more of: deserialize data on the data plane, recover a clock signal, process control characters.

7. The bridge device as claimed in any of the claims 2 to 6, wherein the state machine is coupled to a first circuit and to a second circuit, wherein each one of the first circuit and the second circuit has a data plane and a control plane.

8. The bridge device as claimed in claim 7, wherein the first circuit is coupled to a first bi-directional physical medium attach; and wherein the second circuit is coupled to a second bi-directional physical medium attach.

9. The bridge device as claimed in claim 7 or 8, wherein the data plane comprises a data physical coding sublayer, coupled to a data physical sublayer.

10. The bridge device as claimed in any of the claims 7 to 9, wherein the control plane comprises a set of buffers, a control physical coding sublayer, a control physical sublayer, a protocol circuit, and a phase locked loops and clock distribution circuit.

11. The bridge device as claimed in any of the preceding claims, comprising a translation circuit configured to perform network address translation.

12. The bridge device as claimed in any of the preceding claims, wherein the bridge device is configured to receive networkprotocol information and wherein the bridge device is protocol agnostic.

13. The bridge device as claimed in any of the preceding claims, comprising a memory.

14. The bridge device as claimed in any of the preceding claims, wherein the first port circuit is adapted to perform optical to electrical conversion, and wherein the second port circuit is adapted to perform electrical to optical conversion.

15. A bridge device system comprising at least one pair of bridge devices, wherein the said at least one pair of bridge devices comprises first bridge device according to any of the claims 1 to 14 and configured to transmit data in a first direction, and a second bridge device according to any of the claims 1 to 14 configured to transmit data in a second direction.

16. The bridge device system as claimed in claim 15, wherein for each pair, the state machine of the first bridge device is coupled to the state machine of the second bridge device.

17. The bridge device system as claimed in claim 15 or 16, wherein the first bridge device is configured to perform a primary conversion from a first optical standard on a receiving side to a second optical standard on a transmitting side and wherein the second bridge device is configured to perform conversion from the second optical standard on a receiving side to the first optical standard on a transmitting side.

18. The bridge device system as claimed in claim 17, wherein the first bridge device has a first pair of physical medium attachments forperforming the primary conversion; and wherein the second bridge device has a second pair of physical medium attachments for performing the secondary conversion.

19. An optical switching network comprising a plurality of bridge devices as claimed in any of the claims 1 to 14.

20. The optical switching network as claimed in claim 19, comprising a plurality of switch devices, each switch device being associated with a corresponding domain; and wherein a bridge device or a bridge device pair is provided between two different domains.

21. The optical switching network as claimed in claim 20, comprising a plurality of hosts, and wherein a host has a different address in different domains.

22. The optical switching network as claimed in any of the claims 19 to 21, wherein each host comprises at least one network interface controller; a network data plane and a network control plane distributed among the plurality of hosts and the plurality of switch devices forming the network; wherein the said at least one network interface controller is configured to provide network protocol information to the network data plane and control information to the network control plane for routing optical data to be transmitted between a source host and a destination host among the plurality of hosts.

Citation Information

Patent Citations

  • Hierarchy of control in a data center network

    EP2774328B1

  • Optical network system

    US11368768B2

  • Compute nodes within reconfigurable computing clusters

    US20200233718A1

  • Multi-protocol network interface card

    US7742489B2