Packet switched control and circuit switched data apparatus orchestrated by low-latency digital programmable controller

The hybrid network architecture using a central orchestration unit dynamically reconfigures circuit-switched networks for ultra-low latency data transfer, addressing inefficiencies in data center applications by synchronizing devices and optimizing data integrity, reducing latency from milliseconds to nanoseconds.

JP2025106171APending Publication Date: 2025-07-14MORGAN STANLEY SERVICES GROUP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025001972
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-11-26
Filing Date
2025-01-06
Publication Date
2025-07-14

AI Technical Summary

Technical Problem

Existing data center applications face inefficiencies in data transmission due to increased latency and bandwidth requirements as the number of devices grows, with packet-switched networks causing congestion, high costs, and circuit-switched systems requiring costly circuit establishment, leading to unacceptable latency in data transfer.

Method used

A low-latency digital programmable controller orchestrates a hybrid network architecture combining packet-switched and circuit-switched networks, using a central orchestration unit to dynamically reconfigure the circuit-switched fabric for ultra-low latency connections, synchronize devices with a shared clock, and employ pattern matching for delimiter detection and descrambling.

Benefits of technology

The solution significantly reduces latency from milliseconds to nanoseconds, improves scalability, and enhances data integrity by synchronizing devices and eliminating time-consuming synchronization processes, while maintaining compatibility with existing protocols.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106171000001_ABST
    Figure 2025106171000001_ABST
Patent Text Reader

Abstract

To provide data requesting devices and data sending devices.SOLUTION: Each of the devices is configured with at least one request port and at least one response port. A packet switched network device is coupled to the respective request ports of the data requesting devices and the data sending devices. Moreover, a low-latency digital programmable controller configured as a central orchestrating unit is coupled to the packet switch network device. A circuit switched network device is coupled to the central orchestrating unit and coupled to at least the respective response ports of the plurality of data requesting devices and the plurality of data sending devices, wherein the circuit switched network device is configured to receive a data request from one of the plurality of data requesting devices for data from one of the plurality of data sending devices, and send the data request to the central orchestrating unit.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - reference to Related Applications This application claims priority to U.S. Provisional Patent Application No. 63 / 616,930, filed on January 2, 2024, the entire content of which is incorporated by reference as if fully set forth herein.

[0002] The present disclosure generally relates to data transmission, and more particularly, to a flexible and scalable latency reduction architecture that connects a requesting device to a transmitting device.

Background Art

[0003] Data center applications support moving data between a requesting device and a responding device. Generally, a requesting device refers to a module that requests access to data (e.g., a CPU or GPU or AI accelerator). A responding device refers to a module that sends data and may include a TPU or AI accelerator, GPU, or a memory bank (coupled with a CPU or otherwise). During operation, generally, requests for data consist of short - length data, and the responses on the responding side are relatively long - length data. A typical network computing application requires that multiple requesting and responding devices can communicate with each other via a network. Over time, the need for additional requesting / responding devices per network often increases. Unfortunately, adding requesting / responding devices requires an increase in data throughput capacity.

[0004] Typical data center applications place request and response side devices on a packet switched network. Unfortunately, packet switched networks can be prone to congestion as the number of connected devices increases, incur data costs per response side device, and this can be super-linear in total response side bandwidth. In operation, packet switching generally supports all devices connected to each other simultaneously, but not all devices need to be transmitting at the same time.

[0005] More devices / more bandwidth means more transistors dedicated to networking, and thus more power, so packet switched networks do not scale efficiently. Additionally, packet switched networks introduce high latency. Circuit switched systems, in contrast, do not have this drawback, but require the establishment of circuits for block-based scrambled digital protocols (64b / 66b and the like (Ethernet, InfiniBand, etc.)), which require costly sinks (above 10 microseconds) in circuit establishment.

[0006] During data transmission between a sending device and a receiving device, the receiving device must "lock" onto the base clock that the sending device is using to send data. By doing so, the receiving device can sample the data, and by doing so, the receiving device can identify the delimiter between data blocks. In an Ethernet-like protocol, this delimiter can be an unscrambled 2-bit sync header (either 01 or 10), which can be different in other network protocols. In any case, the concept is generally the same in the latest L2 protocols, and the receiving device relies on the delimiter to correctly parse individual data blocks. Moreover, the receiving device must synchronize its scrambler state with that of the sending device to correctly descramble the data. Unfortunately, these steps are time-consuming and, in the case of an Ethernet-like protocol network, can exceed 10 microseconds and can occur during each circuit-switched event. This results in unacceptable latency and cancels out other performance benefits.

[0007] The disclosure made herein is presented with respect to these and other considerations. SUMMARY OF THE INVENTION MEANS FOR SOLVING THE PROBLEM

[0008] In one or more implementations, a data apparatus and method are disclosed, which are orchestrated by a low-latency digital programmable controller. In one or more implementations, a plurality of data requester devices and a plurality of data sender devices are provided, and each of the data requester devices and data sender devices is composed of at least one request port and at least one response port. Further, a packet-switching network device is coupled to the respective request ports of at least the plurality of data requester devices and the plurality of data sender devices. Moreover, a low-latency digital programmable controller configured as a central orchestration unit is coupled to the packet-switching network device. A circuit-switching network device is coupled to the central orchestration unit and is coupled to the respective response ports of at least the plurality of data requester devices and the plurality of data sender devices, and the circuit-switching network device is configured to receive a data request from one of the plurality of data requester devices for data from one of the plurality of data sender devices and send the data request to the central orchestration unit. Further, the central orchestration unit is configured to receive a data request from the packet-switching network and process the data request by dynamically reconfiguring the circuit-switching network fabric to create a physical connection between one of the plurality of data requester devices and one of the plurality of data sender devices.

[0009] In one or more implementations of the present disclosure, the central orchestration unit is an ultra-low-latency field-programmable gate array or an application-specific integrated circuit.

[0010] In one or more implementations of the present disclosure, the packet-switching network is an Ethernet switch.

[0011] In one or more implementations of the present disclosure, the data apparatus includes a request plane managed by a central orchestration unit and a data response plane.

[0012] In one or more implementations of the present disclosure, a circuit - switched network fabric is reconfigured based on predetermined configuration information representing the circuit - switched network.

[0013] In one or more implementations of the present disclosure, one of a plurality of data requester devices seeking data from one of a plurality of data sender devices is mesosynchronized prior to one of the plurality of data sender devices transmitting data to one of the plurality of data requester devices.

[0014] In one or more implementations of the present disclosure, one of a plurality of data sender devices transmits a data pattern along with the requested data, and further, the data pattern is known in advance by one of the plurality of data requester devices, and further, one of the plurality of data requester devices uses the data pattern to identify delimiters in the transmitted data.

[0015] In one or more implementations of the present disclosure, data transmitted by one of a plurality of data sender devices is scrambled, and further, one of the plurality of data requester devices knows the scrambler state and the scrambler type of one of the plurality of data sender devices, and further, one of the plurality of data requester devices uses the scrambler state and the scrambler type to descramble the data.

[0016] In one or more implementations of the present disclosure, at least one other low - latency digital programmable controller is configured as at least one second central orchestration unit coupled to a packet - switched network device.

[0017] Aspects of the present disclosure will be more readily understood by considering the detailed description of its various embodiments, which are described below in conjunction with the accompanying drawings.

Brief Description of the Drawings

[0018]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

DETAILED DESCRIPTION OF THE INVENTION

[0019] The present disclosure provides a system and method for packet switching control, including a circuit-switched data device programmed by a low-latency digital programmable controller core. Referring now to FIG. 1, an exemplary programmable architecture 100 according to an exemplary implementation of the present disclosure is shown. In the example shown in FIG. 1, a central orchestration unit 102 is electrically coupled to a circuit-switched network 104 and a packet-switched network 106. In the implementation shown in FIG. 1, the central orchestration unit 102 is an ultra-low latency field programmable gate array (FPGA). In one or more implementations of the present disclosure, the central orchestration unit 102 may be an application specific integrated circuit (ASIC). As shown in FIG. 1, the architecture 100 includes a request plane (managed by the central orchestration unit 102) (low latency but low bandwidth) and a data response plane (low latency and high bandwidth). Also, in the implementation shown in FIG. 1, the circuit-switched network 104 is configured as a crossbar switch and the packet-switched network 106 is configured as a layer 1.5 switch of an Ethernet-like protocol. More specifically, the orchestration plane of the central orchestration unit 102 is coupled to the crossbar orchestrator of the crossbar switch.

[0020] Continuing with reference to FIG. 1, a plurality of requester / responder devices 108 (Device A, Device B, Device C, and Devices 1, 2, 3) are shown, each coupled to the circuit-switched network 104 and the packet-switched network 106. More specifically, each request port of the devices 108 is coupled to the packet-switched network 106 and the response port is coupled to the circuit-switched network 104. Each device 108 has one or more ports (response ports) dedicated to sending and receiving data that is high bandwidth, low latency, and large amounts of data. One or more ports are dedicated to sending and receiving control data that is low bandwidth, ultra-low latency, and small amounts of data (request ports).

[0021] During operation, the programmable architecture 100 can, as a function of the central orchestration unit 102, dynamically reconfigure the fabric of the circuit - switched network 104 to create a physical connection between two of the devices 108. The central orchestration unit 102 can be a programmable sub - 10 - nanosecond circuit (e.g., an FPGA) and can be connected to the packet - switched network 106 (e.g., operating as a control fabric) to receive all requests sent to the packet - switched network 106. The central orchestration unit 102 can communicate with the request ports of all devices. The devices 108 cannot see the ports of other devices 108 through packet switching. The central orchestration unit 102 can, for example, configure the mapping of the cross - bar fabric programmatically in less than 50 nanoseconds. Once connected, data can move between two devices 108 (e.g., device A and device 1) through the cross - bar switch via the connection points dynamically established by the central orchestration unit 102.

[0022] In one or more implementations of the present disclosure, a central orchestration unit 102 (e.g., an FPGA) is composed of information representing the topology of a circuit-switched network 104 (e.g., a crossbar). In response to receiving a connection request via a packet-switched network 106, the central orchestration unit 102 effectively programs the circuit-switched network 104 to establish a data connection between a device 108 (e.g., device A) and another device 108 (e.g., device 1). In short, a request goes out via the packet-switched network 106, and the central orchestration unit 102 creates a new connection via the circuit-switched network 104. Those skilled in the art will appreciate the improvement in the efficiency of orchestrating connectivity with respect to the scalability and scalability of the architecture 100 as well as contention and errors, and due to dynamic reconfigurable connections when each of the plurality of devices 108 communicates. The present disclosure overcomes the limitations of static physical connections set in the switch, as well as the requirements for more sophisticated computing capabilities (e.g., beyond the computing capabilities of a simple crossbar) in the circuit-switched network 104. The present disclosure effectively utilizes the computing capabilities of the central orchestration unit 102 to support replacing an electrical crossbar with an optical crossbar and scaling up the number of devices 108 that can be interconnected. Different from known systems, the present disclosure supports coding connections on a sub-microsecond time scale. This can be accomplished by moving the requirements for data connectivity from the requesting device 108 via the packet-switched network 106 and simultaneously reconfiguring the circuits in the circuit-switched network to transfer data between the devices 108. As a result, the performance improves from milliseconds (and above) to nanoseconds, an improvement on the order of about six digits. The packet-switched network can operate with very low latency, but in doing so, it has low bandwidth limitations. During operation, for example, the packet-switched network may start and end transmitting data in 30 nanoseconds, but it cannot move 400 gigabits of data per second.Accordingly, the present disclosure supports the separation of data transmission via the data plane and requests via the orchestration plane.

[0023] FIG. 2 shows a flowchart illustrating steps executed in a programmable architecture 100 for data transmission according to an exemplary implementation of the present disclosure. At step 202, the process begins, after which a device 108 ( "Device 1") sends a request for data from another device 108 ( "Device A") via a request port (step 204). Thereafter, the request of Device 1 is forwarded to the central orchestration unit 102 via the packet switched network 106 (step 206). The central orchestration unit 102 forwards the request to Device A (step 208) in response to connecting the response port of Device 1 to the response port of Device A on the circuit switched network 104 (step 210). For example, the central orchestration unit 102 forwards this request to the response device through the packet switched network 106. At step 212, a determination is made as to whether the mapping of the circuit switched network 104 is complete. If not, the process loops until the mapping is complete.

[0024] The central compilation unit 102 (e.g., FPGA) and the crossbar switch may be deterministic. Thus, when the central compilation unit 102 receives a request from a packet - switched network, it can calculate a fixed amount of time between the receipt of that request and the completion of the switch's remapping. Thereafter, a simple fixed - time delay can be implemented on all devices. However, unfortunately, neither the requester nor the response device can determine the exact time at which the central compilation unit 102 receives the request. Moreover, the remapping time need not always be deterministic. The mapping of physically heterogeneous channels may take longer than that of more closely - located channels, and clock skew may cause delays that increment in multiples of the clock period (a slightly delayed clock edge disturbs the setup and hold times of the incoming signal, so the signal is not latched until the next clock edge). Thus, a simple fixed - time approach may not be feasible in a particular implementation form.

[0025] Accordingly, and as referred to herein, the present disclosure provides an improved methodology by using the connection between the requester and the response device 108. More specifically, after the connection is made, initial sink / pattern matching can be performed, where the response device 108 sends an initial known pattern and the requester device 108 returns the same pattern. When the response device 108 sees the same pattern that it previously sent on the receiving side of its response port, it knows that the connection is established. Those skilled in the art will recognize that an Ethernet port may include a device composed of an RX and TX pair. Thus, each response port connected to the circuit - switched network 104 (e.g., crossbar switch) includes RX and TX. Thus, during operation, the TX side of the response port on device A is connected to the RX side of the response port on device 1.

[0026] If the determination in step 212 is affirmative, the process branches to step 214, and device A transmits data to device 1 through the circuit - switched network 104 via each response port of the device. Thereafter, in step 216, it is determined whether the data transmission has been completed. If not, the process loops until the transmission is completed. If the determination in step 216 is affirmative, the process branches to step 218, and device A transmits a data completion confirmation from each of its response ports. Thereafter, the process continues to step 220, where device A transmits the completion confirmation, which is forwarded to the request plane of the circuit - switched network 104 via the packet - switched network 106. Thereafter, the central orchestration unit 102 re - configures the resources on the circuit - switched network 104 to a blank mapping. The features and details of the steps related to FIG. 2 will be described in more detail below.

[0027] It should be recognized here that, despite the performance benefits of separating the request and data layers to establish a connection between two devices 108 (e.g., device A and device 1 as shown and described herein), simply connecting the two devices 108 will not normally result in successful data transfer and reception. Device 108 requires information such as the specific frequency at which the signal operates and encoded information.

[0028] During operation, the clock signal is expected to be the same for devices on the network. In the case of a 10G Ethernet network, for example, the line clock signal (e.g., the base clock signal for data transmission) is 10.3125 GHz. At least some of the devices communicating on the network can transfer via the 10.3125 GHz signal by generating a relatively slow (e.g., 156.25 MHz) signal, and then use a phase-locked loop to generate a 10.3125 GHz signal from the relatively slow signal. For example, considering 66 bits per frame, 156.25Mhz × 66 = 10.3125 GHz. Thus, the clock signal frequency for each frame may be 156.25 MHz, and the clock frequency for each bit is 10.3125 Ghz.

[0029] It should be recognized here that all devices could be configured with a clock signal frequency of 10.3125 GHz, but different devices may use their respective reference clocks to generate relatively fast signals. Thus, all devices may nominally operate at 10.3125 GHz, but minor differences in the tolerance between signals, for example due to their respective reference clocks, can cause data integrity problems.

[0030] At least in a typical 10GBASE-KR Ethernet connection, the present disclosure solves such data integrity problems. In one or more implementations, a transmitting device sends data to a receiving device (RX), and following reception, the receiving device extracts a clock signal from the data received from the TX. Once extracted, the RX uses the TX's clock signal to generate a relatively fast transmission signal. Thus, for example, in the case of a single Ethernet connection having two connected devices, one device with TX and RX and the other device with RX and TX, there can be two clock domains used to generate relatively fast signals. Of course, other design implementations for dealing with the matters related to data integrity described above will be recognized by those skilled in the art. For example, in one or more implementations of the present disclosure, a device can use a clock signal extracted from a received signal to generate a relatively fast signal for transmitted data in order to "time" the transmitted data. This approach can bring all devices into a single clock domain, enabling more convenient handling. Such cases can depend on a particular Ethernet implementation. For example, a device may be designated as a clock master, and the clock signal used for the device's TX data may be used as the RX and TX clocks in the receiving device. In an alternative implementation, another external device can provide clock references to both of the connected devices (e.g., one device's TX and RX, and the other device's RX and TX). Although seemingly a straightforward solution for improving data integrity, the latter implementation can be difficult to scale when many devices are sending and receiving data. For example, the relatively long wires required for the clock signal can cause clock delay and clock skew, and as a result, the clock will "appear" differently in different parts of the circuit. Thus, the present disclosure provides an improved mechanism for connectivity and data transfer.

[0031] In an exemplary operation, the central orchestration unit 102 sends a request to the request port of the data sending device 108 and requests that data be sent at the response port. The sending device 108 transmits a known preamble, such as scrambled / gearbox-in or otherwise, that is placed in front of the data being sent, which is used to lock in the state of the scrambler in the requesting device 108. The length of the preamble is known in advance and need only be sufficient to recover the scrambler state, assuming that the phase-locked loop ("PLL") is synchronized and that full resynchronization is not required, which is an expected condition when the PLL is mesosynchronized and kept within its adjustment range. When the sending device 108 has completed sending data on its response port, the requesting device 108 sends an acknowledgment on its request port, thereby notifying the central orchestration unit 102 that the data transmission has been completed. If the requesting device 108 loses either the PLL or the scrambler sink, the requesting device 108 notifies the central orchestration unit 102 that its transmission has been completed and waits internally until it can re-establish the PLL lock. Thereafter, the requesting device 108 may resubmit its request for the data as necessary (loop back to step 204 in FIG. 2). Thereafter, the central orchestration unit 102 may freely remap any resources consumed by the circuit to any state.

[0032] Accordingly, in one or more implementations of the present disclosure, each device 108 has at least two ports. One or more of the ports may be dedicated to sending and receiving data that may be high bandwidth, low latency, high power, and large amounts of data (response ports). One or more of the ports may be dedicated to sending and receiving control data that is low bandwidth, ultra-low latency, and small amounts of data (request ports). The response ports of all devices 108 are connected within a crossbar switch fabric (e.g., the circuit-switched network 104), and the request ports of all devices 108 are connected to an Ethernet-like protocol switching fabric (i.e., the packet-switched network 106).

[0033] According to the present disclosure, the central orchestration unit 102 dynamically reconfigures the circuit - switched fabric via a programmable sub - 10 - nanosecond circuit (e.g., FPGA). The central orchestration unit 102 is connected to a packet - switched network 106 (control fabric) and receives all requests sent into the control fabric. The central orchestration unit 102 can communicate with the request ports of all devices, and devices cannot see the ports of other devices through packet switching. During operation, the central orchestration unit 102 can configure the mapping of the cross - bar fabric programmatically, for example, in less than 50 nanoseconds.

[0034] When a circuit is established, a phase - locked loop ( "PLL") and a lock including for the scrambled state are established. During operation, the PLL across the data plane is meso - synchronized by any continuous known pattern transmitted when the source device 108 is not receiving data and is thus connected to a clock source on the circuit cross - bar, or by a separate clock source distribution mechanism. For example, the circuit - switched fabric may be filled with a known pattern, thereby feeding a constant signal to the devices connected to the circuit - switched fabric, thereby enabling the devices to lock onto the same clock signal. This eliminates the need to lock in the PLL from a cold state during circuit establishment, which can take much longer than 1 microsecond. During operation, the source device 108 can correctly establish the phase by testing a few phases. The scrambler can achieve fast synchronization by causing each response side to send a known pattern that locks within 1 - 16 bytes depending on the size of the network and the particular protocol when received, so it does not need to be synchronized.

[0035] Regarding unlock and lock latency, the present disclosure ensures that the receiving device 108 “locks” onto the base clock used by the sending device 108 to send data. The receiving device 108 does this in order to be able to sample the data. Additionally, the receiving device 108 locates the delimiter between two or more data blocks. In the case of an Ethernet-like protocol, the delimiter may be an unscrambled 2-bit sync header (either 01 or 10). Other network protocols may use different delimiters, but it will be recognized by those skilled in the art that the concept is shared across most of the current L2 protocols. The receiving device 108 uses the delimiter to correctly parse the separate data blocks. Further, the receiving device 108 synchronizes its scrambler state with the scrambler state of the sending device 108 that is used to correctly descramble the data.

[0036] It should be recognized here that these steps are time-consuming and in a network with a typical Ethernet-like protocol, may take more than 10 microseconds to complete. Such a delay would be unacceptable by negating any performance benefits of this architecture 100 shown and described herein that would occur during each circuit switching event.

[0037] Accordingly, the present disclosure addresses and solves the latency resulting from unlock and lock operations. One simple solution is to ensure that all devices 108 are given the same clock. This can be done, for example, by transmitting a synchronization signal to all devices 108 on the circuit - switched network 104 when the devices 108 do not necessarily communicate actively with each other. Effectively, the synchronization signal feed provides all devices on the crossbar with an opportunity to sink and maintain their respective PLLs with each other. In an alternative implementation, a separate set of connections can be used as a clock tree for distributing the shared clock. Any of the options ensures that all devices 108 operate from the same clock, and thus, the PLL locking time can be limited to 10 - 20 cycles (900 - 1800 ps for 10GBASE - KR or less than 900 - 1800 ps for faster protocols).

[0038] In addition to synchronizing the PLLs, the present disclosure addresses the time - consuming process of detecting block delimiters that are not normally distinguishable from standard data bits in the raw data stream. The present disclosure overcomes the known latency for detecting delimiters that normally results from searching for multiple data packets for a header packet that is constant across multiple packets. Instead, the present disclosure uses a kind of pattern matching, where the transmitting device 108 adds a short known pattern to any data it transmits in other ways, and the receiving device 108 simply locates the known pattern and uses it to identify the delimiter.

[0039] It should be recognized here that the data is usually not transmitted in its raw format, but is ordinarily scrambled. In such cases, adding a known pattern to the data being transmitted may not be effective. It should be further recognized that most scrambling methods are based on a linear feedback shift register with an initial state. When the scrambled data is received and the scrambler type, as well as the initial state of the scrambler, is known, the data can be easily descrambled and the scrambling process is reversible.

[0040] Accordingly, the present disclosure further includes, for each sending device 108, defining a specific scrambler state shared with the receiving device 108 prior to data transmission. For example, in one or more implementations, a data table can be shared with the device at the start of transmission. By referring to the table, each device can identify the initial scrambler state of each other device. Alternatively, a single initial scrambler state can be shared and may be known for all devices.

[0041] The receiving device 108 knows the sending device 108 and the receiving device 108 is synchronized (synchronizes) with the sending device 108. Further, the receiving device 108 knows the initial scrambler initial state of the sending device 108 and knows the initial pattern that the sending device 108 is sending with the data. These enable the receiving device to avoid a time-consuming block synchronization process for a much faster pattern matching process.

[0042] Accordingly, and as shown and described herein, the present disclosure provides an improved packet switched control, circuit switched data device programmed by a low latency digital programmable controller. The architectures shown and described herein support improved flexibility and performance, such as by supporting a crossbar switch that may be electrical or optical, with the optical being more scalable due to lower power and increased channel counts. Further, while many of the implementations and examples shown and described herein consider a single central orchestration unit 102, the present disclosure is not so limited and can support multiple central orchestration units 102 that together form a tree of central orchestration units 102, similar to a switching network, where the tree is either connected by layer 1 or similar switches or is vertical by using some ports on each central orchestration unit 102 as uplinks. Moreover, routing algorithms such as standard industry packet routing algorithms, including packet header based dynamic switching (such as Ethernet-like protocols) or prefix based fixed switching such as InfiniBand, can be used between the multiple central orchestration units 102.

[0043] Further, and as shown and described herein, device 108 is meso-synchronous, whereupon, when a circuit is created, device 108 needs to sync up within a few nanoseconds. This can be accomplished by eliminating the need for full PLL / scrrambler state recovery. Instead, a lock can be established on the most recently connected circuit. In one or more implementations, remapping of the circuit switched network 104 can be blocked in response to possible new requests based on each data transmission.

[0044] As mentioned herein, the start of the connection devices may be synchronized to the same clock when interconnected in a circuit switched network 104 (e.g., a crossbar switch). In operation, this may occur at the start of the connection, where one device extracts the clock from the data sent from the other device and then proceeds using that clock. The time-wasting nature of this approach is overcome by configuring all response ports connected to the circuit switched network 104 crossbar to be meso-synchronized. The present disclosure introduces two solutions to such time problems, namely, that all devices may be directly connected to some external clock reference, or that the device response ports may be directly fed the same clock via the circuit switched network 104 (e.g., a crossbar switch). In one or more implementations, the latter approach is preferred as the crossbar switch is effectively "filled" with the clock signal from the central orchestration unit 102. This clock provides the basic clock signal for the intended data protocol. For example, in the case of 10GBASE-R, the clock signal frequency is 10.3125 GHz.

[0045] In one or more implementations, the default (blank) mapping of the circuit switched network 104 (e.g., a crossbar switch) connects all of the RX sides of the response ports for device 108 to the clock signal from the central orchestration unit 102. FIG. 3 shows an exemplary configuration of the circuit switched network 104 (e.g., a crossbar switch) for such connectivity. As shown in FIG. 3, the dotted lines represent lines carrying the clock from the central orchestration unit 102. As used herein, this clock may generally be referred to as the "COU sync clock".

[0046] Sometimes, the two devices 108 need to be connected to each other, as a result of which the clock sink may be disconnected, and by doing so, it should be recognized here that the lock between those two devices is lost. To address this problem, the present disclosure addresses it by defining a first data type that the response device transmits as a training sequence consisting of 1s and 0s via its TX channel. The training sequence can be generated within the clock domain of the COU sink clock received on the RX channel. This operation is exactly the same as the clock itself. In such a case, the training sequence is meso-synchronized with the COU sink clock. In other words, the training sequence will have the same frequency as the COU sink clock, but does not have to be in the same phase.

[0047] The present disclosure addresses the concern that the above-described switchover is hitless, that is, the switch from the COU sink clock to the training sequence TX channel creates no glitches or delays. It should be recognized here that some interruption at the very moment of the switchover can be envisioned, for example, due to the clock being out of phase. Such glitches or interruptions are likely to be of little importance because the receiver's PLL is already synchronized with the COU sink clock, such as having the same frequency as the training sequence, and the PLL is minimally affected. FIG. 4 shows an exemplary switchover event between two devices according to an exemplary implementation of the present disclosure. Further, FIG. 5 shows exemplary glitches resulting from a switch from one device to another with different clock phases.

[0048] As referred to herein, the present disclosure enables a receiving device 108 to identify delimiters between two or more data blocks. In one or more implementations of the present disclosure, 10GBASE-R Ethernet using 64 / 66b encoding is provided. As is known in the art, 64 / 66b encoding involves sending data within 66-bit frames, where each frame includes two sync header bits and 64 scrambled data bits. The two sync header bits are defined as either 0 1 or 1 0. In other words, the sync header bits include a bit transition (rising edge or falling edge), but the sync header bits do not support values of 1 1 or 0 0 as opposed to the 64 scrambled data bits. The 64 scrambled data bits may appear random, but those skilled in the art will recognize that such scrambling is not a random process and that a linear feedback shift register or other feature may be used to create pseudo-random data. FIG. 6 shows a periodic representation of 64 / 66b encoding according to an exemplary implementation of the present disclosure. The exemplary implementations shown and described herein consider 10GBASE-R Ethernet, but it should be understood that other encoding styles, such as 8 / 10b, are supported by the present disclosure when the logical principles are applicable.

[0049] Continuing with the example shown in FIG. 6, a 66-bit frame can include positions for bit transitions, and the remaining positions include data that appears to be randomized. For example, 100 66-bit frames can be aligned based on the positions for bit transitions, thereby "defining" the boundaries between each 66-bit block, differentiating, and thus enabling the ability to determine where to extract each 64-bit data.

[0050] The process of determining the boundaries between blocks is generally referred to herein as block sync, which occurs once at the start of an Ethernet connection. Those skilled in the art will recognize that there are various block sync options. One is for the FPGA transceiver to randomly select two adjacent bits in the stream and then implement a process of checking transitions over multiple frames. If no transition is detected, the FPGA "shifts" forward by only one bit and checks the next two adjacent bits. The FPGA repeats this process until two adjacent bits with a transition are detected. The FPGA can then identify the location of the delimiter (e.g., relative to the rest of the transceiver).

[0051] Unfortunately, the block synchronization process can be time-consuming. For example, false negatives can be detected, where the process locks onto that bit position for some time until no transition is found at the scrambled bit position and then moves to the next position. This time cost is usually only incurred at the start of the connection, so it is not a problem, and in a normal network, the connection is not disconnected and is often reconnected. Embodiments of the present disclosure consider network connectivity involving devices that frequently connect and disconnect. The features shown and described herein accelerate that process, more specifically, the block sync process, by transmitting a specific pattern that the block sync lock quickly identifies.

[0052] In one example, a particular pattern is a 66-bit block with a sync header, where the 64 payload (data) bits consist of either all 0s or all 1s. This can be envisioned as a frame consisting essentially of only the sync header. Using such a pattern, even when the block sync circuit selects a bit position immediately following the sync header, there is only one clock cycle on each remaining bit until the sync header is located. From this point on, the transceiver is locked onto the delimiter and the data can be sent normally. In this example, in such a case, it should be understood that the response device sends at least 66 copies of the "sync header only" frame. At 10GBASE-R speed, each bit is 96.96 ps. Therefore, it takes 6.399 ns to send a 66-bit frame and 422357 ps (422.4 ns) to send 66 such frames.

[0053] In one or more implementations, the present disclosure can improve this performance. For example, after the initial clock, a known pattern that the response side sends first can be designed, and this clock can be provided per destination or globally. In any case, the known pattern is preferably predefined and shared among all devices. Each receiving device can check when this pattern is received and then run a simple pattern match to lock onto the last bit of this pattern as a delimiter. Supporting the implementation in this way may include modifying the block sink circuit in the transceiver to recognize the known pattern. However, the use of such a simple 66-bit pattern can further reduce the lock time, for example, from 422.4 ns to 6.4 ns. Further, the results are backward compatible with existing transceivers (e.g., while operating as a receiver), and thus do not require any changes to the Ethernet protocol. Additionally, in one or more implementations of the present disclosure, the training sequence can be a valid sequence. In such a case, the transceiver can lock on, for example, after 66 frames, and regardless, the receiver operates from the clock, and the circuit can more quickly determine where the fixed delimiter is between each packet.

[0054] After an Ethernet connection is made between two devices, it may take some time for the scrambler states to match (e.g., by a linear feedback shift register), which enables the data to be descrambled, as would be recognized by those skilled in the art. While the transmitting scrambler linear feedback shift register and the receiving descrambler linear feedback shift register are synchronized, the scrambled data from the scrambling linear feedback shift register is sent until all the stored bits are overwritten. This process is relatively fast because linear feedback shift registers typically do not store many bits, but it can take some time and cost for many instances. The present disclosure overcomes this problem with a solution that defines all nodes within a circuit-switched network to have the same scrambler state at the instant of connection. Since scrambling is not used for security purposes, those skilled in the art will recognize that security is not compromised and speed is improved as long as all devices pre-agree to use a specific known scrambler state and each scrambler is loaded / reloaded at the instant the connection is made. Orchestrating the sharing of scrambler states can be done once at startup by the packet-switched network and is not overly complex to achieve, as would be understood by those skilled in the art. This option eliminates the need for a scrambler sink, thereby saving time, e.g., by 5.7 ns.

[0055] FIG. 7 is a schematic diagram showing transmitted data when a connection is made on a circuit-switched network 104. The access time described according to the present disclosure is significantly improved. The typical access time in the case of DRAM is about 100 cycles, slightly exceeding 40 ns. This means that within a single rack device, it can take some time for the CPU to access data from DRAM. Embodiments of the present disclosure support one rack device accessing another rack device within a similar time period. This represents a significant improvement.

[0056] Moreover, the present disclosure addresses the possible trade - offs between one - to - one blocking, as opposed to the central orchestration unit 102 stacking requests prior to parallel mapping. In fact, parallel mapping may allow multiple devices to use the central orchestration unit 102, but the central orchestration unit 102 may need to delay processing until a sufficient number of data transmission requests are received. The present disclosure solves this by utilizing a crossbar for the circuit - switched network 104, which can be configured per channel.

[0057] Any of the features shown and described herein may, in a corresponding implementation, be reduced to a non - transitory computer - readable medium (such as a CRM like a disk drive or flash drive) storing computer instructions, which, when executed by a processing circuit, cause the processing circuit to practice an automated process for implementing each method.

[0058] It should be further understood that like or similar numbers in the drawings represent like or similar elements throughout several figures, and that all components or steps described and shown with reference to the drawings are not necessarily required for all embodiments or configurations.

[0059] The terms used herein are for the purpose of describing particular embodiments only and are not intended to be limiting of the present disclosure. As used herein, the singular forms “a,” “an,” and “the” are to be construed to include the plural forms as well, unless the context clearly dictates otherwise. Further, the terms “comprises” and / or “comprising,” when used herein, specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0060] Any directional terms are used herein only for purposes of agreement and reference and are not intended to be limiting. However, it should be recognized that these terms may be used with reference to the viewer. Accordingly, no limitation is implied or should be inferred. Further, the use of ordinal numbers (e.g., first, second, third) is for purposes of distinction and not for counting. For example, the use of "third" does not imply the existence of a corresponding "first" or "second". Also, the style and terminology used in this specification are for descriptive purposes and should not be regarded as limiting. The use of "including", "comprising", "having", "containing", "accompanying", and variations thereof in this specification is intended to cover the items listed thereafter and their equivalents as well as additional items.

[0061] The above-described subject matter is presented by way of example only and is not intended to be limiting. Various modifications and changes can be made to the subject matter described in this specification without departing from the true spirit and scope of the invention as encompassed by this disclosure, and without following the illustrated and described exemplary embodiments and the scope of their application. The spirit and scope of the invention are defined by the set of recitations in the appended claims and by structures and functions or steps that are equivalent to these recitations.

Description of Reference Numerals

[0062] 100 Programmable architecture, architecture 102 Central composition unit 104 Circuit-switched network, circuit-switched network 106 Packet-switched network 108 Requestor / responder device, device, data sending device, sending device, receiving device

Claims

1. A data device programmed by a low-latency digital programmable controller, a plurality of data requester devices and a plurality of data sender devices, each of the data requester devices and the data sender devices being configured using at least one request port and at least one response port, a packet switching network device coupled to at least the respective request ports of the plurality of data requester devices and the plurality of data sender devices, a low-latency digital programmable controller configured as a central programming unit coupled to the packet switching network device, a circuit switching network device coupled to the central programming unit and at least to the respective response ports of the plurality of data requester devices and the plurality of data sender devices, wherein the circuit switching network device, receives a data request from one of the plurality of data requester devices for data from one of the plurality of data sender devices, and is configured to send the data request to the central programming unit, and further, the central programming unit, receives the data request from the packet switching network, and is configured to process the data request by dynamically reconfiguring the fabric of the circuit switching network to create a physical connection between the one of the plurality of data requester devices and the one of the plurality of data sender devices. An apparatus.

2. The apparatus according to claim 1, wherein the central programming unit is an ultra-low-latency field programmable gate array or an application-specific integrated circuit.

3. The apparatus according to claim 1, wherein the packet switching network is an Ethernet switch.

4. The apparatus according to claim 1, wherein the data device includes a request plane managed by the central programming unit and a data response plane.

5. The apparatus according to claim 1, wherein the fabric of the circuit switching network is reconfigured based on predetermined configuration information representing the circuit switching network.

6. The one of the plurality of data requestor devices that requests data from one of the plurality of data sender devices is mesosynchronized prior to the one of the plurality of data sender devices transmitting data to the one of the plurality of data requestor devices, the apparatus of claim 1.

7. The one of the plurality of data sender devices transmits a data pattern along with the requested data, and further, the data pattern is known in advance by the one of the plurality of data requestor devices, and further, the one of the plurality of data requestor devices uses the data pattern to identify a delimiter in the transmitted data, the apparatus of claim 1.

8. The data transmitted by the one of the plurality of data sender devices is scrambled, and further, the one of the plurality of data requestor devices knows the scrambler state and the scrambler type of the one of the plurality of data sender devices, and further, the one of the plurality of data requestor devices uses the scrambler state and the scrambler type to descramble the data, the apparatus of claim 7.

9. The apparatus of claim 1, further comprising at least one other low-latency digital programmable controller configured as at least one second central orchestration unit coupled to the packet switching network device.

10. Configuring each of the plurality of data requestor devices and the plurality of data sender devices using at least one request port and at least one response port; Coupling a packet switching network device to at least the respective request ports of the plurality of data requestor devices and the plurality of data sender devices; Configuring a low-latency digital programmable controller as a central orchestration unit coupled to the packet switching network device; A data orchestration method including coupling a circuit switching network device to the central orchestration unit and to at least the respective response ports of the plurality of data requestor devices and the plurality of data sender devices, wherein the circuit switching network device is Receiving a data request from one of the plurality of data requesting devices for data from one of the plurality of data sending devices; Configured to send the data request to the central orchestration unit; Furthermore, the central orchestration unit: Receives the data request from the packet switched network; Processes the data request by dynamically reconfiguring the fabric of the circuit switched network to create a physical connection between the one of the plurality of data requesting devices and the one of the plurality of data sending devices. **Claim 11** The method according to claim 10, wherein the central orchestration unit is an ultra-low latency field programmable gate array or an application specific integrated circuit. **Claim 12** The method according to claim 10, wherein the packet switched network is an Ethernet switch. **Claim 13** The method according to claim 10, wherein the data device includes a request plane managed by the central orchestration unit and a data response plane. **Claim 14** The method according to claim 10, wherein the fabric of the circuit switched network is reconfigured based on predetermined configuration information representing the circuit switched network. **Claim 15** The one of the plurality of data requesting devices that requests data from one of the plurality of data sending devices is mesosynchronized prior to the one of the plurality of data sending devices transmitting data to the one of the plurality of data requesting devices. **Claim 16** The one of the plurality of data sending devices transmits a data pattern along with the requested data, and the data pattern is known in advance by the one of the plurality of data requesting devices, and the one of the plurality of data requesting devices uses the data pattern to identify a delimiter in the transmitted data. **Claim 17** The data transmitted by the one of the plurality of data sending devices is scrambled, and One of the plurality of data requesting devices knows the scrambler state and the scrambler type of one of the plurality of data sending devices, and further, one of the plurality of data requesting devices uses the scrambler state and the scrambler type to descramble the data. The method according to claim 16.

18. The method according to claim 10, further comprising the step of configuring at least one other low latency digital programmable controller as at least one second central orchestration unit coupled to the packet switching network device.