Systems and methods for nodes communicating using time-synchronized transport layers

By employing the Time Synchronization Transport Layer (TSL) protocol and Time Division Multiplexing (TDM) technology in distributed computing networks, node clocks are synchronized and static scheduling is performed, solving the problems of network error susceptibility and reduced throughput, and achieving efficient and reliable data transmission.

CN114641952BActive Publication Date: 2025-10-24MICROSOFT TECHNOLOGY LICENSING LLC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080077417.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-11-06
Filing Date
2020-10-27
Publication Date
2025-10-24
Estimated Expiration
2040-10-27

AI Technical Summary

Technical Problem

In distributed computing networks, existing technologies may reduce throughput while ensuring message exchange integrity, and the network is prone to errors, leading to packet retransmission and performance degradation.

Method used

The Transport Layer Time Synchronization (TSL) protocol is adopted to synchronize the clocks of all participating nodes in the network and use Time Division Multiplexing (TDM) technology to achieve static scheduling of data transmission, avoid link congestion and packet retransmission, and ensure the reliability and efficiency of data transmission.

Benefits of technology

It improves network link utilization, reduces the number of ACKs and data retransmissions, and enhances data transmission reliability and throughput, making it suitable for high-performance backend network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114641952B_ABST
    Figure CN114641952B_ABST
Patent Text Reader

Abstract

Systems and methods are provided for providing message transport between nodes (e.g., acceleration components configurable to accelerate services) using a time-synchronous transport layer (TSL) protocol. An example method in a network including at least a first node, a second node, and a third node includes each of the at least first node, second node, and third node synchronizing a respective clock to a common clock. The method further includes each of the at least first node, second node, and third node scheduling data transmission in the network in a manner such that at a particular time with reference to the common clock, each of the at least first node, second node, and third node is scheduled to receive data from only one of the first node, second node, or third node.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Users are increasingly accessing applications provided via computing, networking, and storage resources located in data centers. These applications run in a distributed computing environment, sometimes referred to as a cloud computing environment. Computer servers in a data center are interconnected via a network, and thus applications running on the computer servers can communicate with each other via the network. These servers can exchange messages with each other using various protocols. Due to the error-prone nature of distributed computing networks, servers can implement retransmission of packets and other schemes to ensure the integrity of message exchanges. However, adding such techniques can degrade throughput.

[0002] Thus, there is a need for methods and systems that mitigate some of these issues. SUMMARY

[0003] In one example, the disclosure relates to a method in a network comprising at least a first node, a second node, and a third node. The method can include each of the at least first node, second node, and third node synchronizing a respective clock to a common clock. The method can further include each of the at least first node, second node, and third node scheduling data transmissions in the network in a manner such that at a particular time with reference to the common clock, each of the at least first node, second node, and third node is scheduled to receive data from only one of the first node, second node, or third node.

[0004] In another example, the disclosure relates to a system comprising a network configured to interconnect a plurality of acceleration components. The system can further include the plurality of acceleration components configurable to accelerate at least one service, wherein each of the plurality of acceleration components is configured to synchronize a respective clock to a common clock, the common clock being associated with an acceleration component selected from among the plurality of acceleration components, and wherein each of the plurality of acceleration components is configured to transmit data in the network in a manner such that at a particular time with reference to the common clock, each of the plurality of acceleration components is scheduled to receive data from only one of the plurality of acceleration components.

[0005] In yet another example, the present disclosure relates to a method in a network comprising at least a first acceleration component, a second acceleration component, and a third acceleration component. The method can include each of the at least first acceleration component, the second acceleration component, and the third acceleration component synchronizing a respective clock with a common clock, the common clock being associated with a selected acceleration component from among the at least first acceleration component, the second acceleration component, and the third acceleration component. The method can further include each of the at least first acceleration component, the second acceleration component, and the third acceleration component scheduling data transmissions in the network in a manner such that at a particular time with reference to the common clock, each of the at least first acceleration component, the second acceleration component, and the third acceleration component is scheduled to receive data from only one of the first acceleration component, the second acceleration component, and the third acceleration component.

[0006] This summary is provided to introduce some concepts of what follows in a simplified form. This summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS

[0007] The present disclosure is illustrated by way of example and not limitation in the figures of which like references indicate similar elements. The elements in the figures are graphically illustrated for simplicity and clarity and are not necessarily drawn to scale.

[0008] Figure 1 A diagram of an acceleration component including a time synchronization transport layer (TSL) component is shown, according to one example;

[0009] Figure 2 A diagram of a TSL component is shown, according to one example;

[0010] Figure 3 A diagram of a node in a network for transmitting messages using a TSL protocol is shown, according to one example;

[0011] Figure 4 A diagram of phases associated with a TSL is shown, according to one example;

[0012] Figure 5 An example of message exchanges between a primary node A, a node B, and another node C during a characterization phase of a TSL is shown;

[0013] Figure 6 An example of message exchanges between a primary node A, a node B, and a node C during a standby phase of a TSL is shown;

[0014] Figure 7An example of message exchange between the main node A, node B and node C during the preparation phase of the TSL is shown;

[0015] Figure 8 An example of message exchange and data transfer between the main node A, node B and node C during the data transfer phase of the TSL is shown;

[0016] Figure 9 An effect of time drift of a local clock relative to a master clock according to one example is shown;

[0017] Figure 10 An example exchange of synchronization messages between the main node A, node B and node C is shown;

[0018] Figure 11 An example of a margin from the perspective of messages sent and received by node B according to one example is shown;

[0019] Figure 12 An example of messages sent and received by node B in a manner that reduces the margin according to one example is shown;

[0020] Figure 13 Another example of message exchange with a margin according to one example is shown;

[0021] Figure 14 Another example of message exchange with an improved margin according to one example is shown;

[0022] Figure 15 An example message exchange including the use of a resilient buffer is shown;

[0023] Figure 16 A flowchart of a method for exchanging messages between nodes using the TSL protocol according to one example is shown; and

[0024] Figure 17 A flowchart of a method for exchanging messages between acceleration components using the TSL protocol according to one example is shown. DETAILED DESCRIPTION

[0025] Examples described in this disclosure relate to methods and systems for providing message transmission between nodes (e.g., nodes including acceleration components that can be configured to accelerate services). Certain examples relate to methods and systems using the Time Synchronization Transport Layer (TSL) protocol. The TSL protocol is designed to minimize the total time to transmit large amounts of data in a dense communication mode (e.g., all-to-all broadcast) over a highly reliable network. Broadly speaking, in order to function properly, examples of the TSL protocol may require data integrity, including almost no packet loss or bit flip errors in the network link, and end-to-end in-order delivery of messages. The TSL protocol works well when certain aspects related to latency, bandwidth, and external packets are met. In terms of latency, the TSL protocol works well if the end-to-end delay between nodes is nearly constant and there is a similar delay between each pair of nodes configured to send / receive messages. With respect to bandwidth, when there is no link congestion, the network switch can support the maximum possible bandwidth on all ports simultaneously. The TSL protocol also works well when there are infrequent non-TSL (or, external) packets. Even if all of these properties are not met, the TSL protocol can be configured to tolerate deviations. Given a network that meets the aforementioned prerequisites, the TSL protocol is able to synchronize all participating nodes (e.g., acceleration components) and run statically scheduled (or dynamically scheduled) data transmission, so that data packets from different senders are not interleaved on the receiver side. This eliminates link congestion and data packet retransmissions, and significantly reduces the number of necessary control messages such as ACKs and data retransmissions.

[0026] The TSL protocol can achieve near-peak link utilization for collective communications by employing time division multiplexing (TDM) technology on top of conventional packet-switched data center networks. In certain examples, fine-grained time synchronization and explicit coordination of all participating endpoints to avoid conflicts during data transmission can improve link utilization of links between nodes. The TSL protocol characterizes network latency at runtime to keep all participating endpoints synchronized with a global clock, and then schedules data transmission across multiple pairs of communicating endpoints to achieve conflict-free full-bandwidth utilization of available network links. The TSL protocol can be best performed in a controlled network environment (e.g., a high-performance backend network) where all endpoints under a TOR or higher-layer T1 / T2 switch set participate in the protocol and coordinate their communications. The TSL protocol is advantageously robust to small delay variations and clock drifts of participating nodes, but in order to minimize hardware overhead and take advantage of the highly reliable and predictable nature of modern data center networks, the TSL protocol can forgo the hardware logic required for retransmission, reordering, and reassembly using simple fail-stop mechanisms or delegating recovery to higher-layer software mechanisms.

[0027] An acceleration component includes, but is not limited to, a hardware component that is configurable (or configured) to perform a function more efficiently than a software implementation running on a general central processing unit (CPU). An acceleration component can include a field-programmable gate array (FPGA), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), an erasable and / or complex programmable logic device (PLD), a programmable array logic (PAL) device, a generic array logic (GAL) device, and a massively parallel processor array (MPPA) device. An image file can be used to configure or reconfigure an acceleration component, such as an FPGA. Information included in the image file can be used to program hardware components of the acceleration component (e.g., logic blocks and reconfigurable interconnects of an FPGA) to implement a desired function. The desired function can be implemented to support any service that can be provided via a combination of compute, networking, and storage resources, such as via a data center or other infrastructure for delivering services.

[0028] The described aspects can also be implemented in a cloud computing environment. Cloud computing can refer to a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be employed in the marketplace to offer ubiquitous and convenient on-demand access to the shared pool of configurable computing resources. The shared pool of configurable computing resources can be rapidly provisioned via virtualization and released with low management effort or service provider interaction, and then scaled accordingly. A cloud computing model can include various features, such as on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service, etc. A cloud computing model can also exhibit various service models, such as Software as a Service ("SaaS"), Platform as a Service ("PaaS"), and Infrastructure as a Service ("IaaS"), etc. The cloud computing model can also employ different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, etc.

[0029] A data center deployment can include a plurality of networked acceleration components (e.g., FPGAs) and a plurality of networked software-implemented host components (e.g., central processing units (CPUs)). A network infrastructure can be shared between the acceleration components and the host components. Each host component can correspond to a server computer that executes machine-readable instructions using one or more central processing units (CPUs). In one example, the instructions can correspond to a service, such as a text / image / video search service, a translation service, or any other service that can be configured to provide useful results to users of devices. Each CPU can execute instructions corresponding to various components (e.g., software modules or libraries) of the service. Each acceleration component can include hardware logic to implement functionality, such as a portion of the service provided by the data center.

[0030] In some environments, software-implemented host components are natively linked to corresponding acceleration components. The acceleration components can communicate with each other via a network protocol. In order to provide reliable service to users of services provided via a data center, any communication mechanism can be required to meet certain performance requirements, including reliability. In certain examples, the present disclosure provides a lightweight transport layer to meet such requirements. In one example, acceleration components can communicate with each other via a network. Each acceleration component can include hardware logic to implement functionality such as a portion of a service provided by a data center. The acceleration components can use several parallel logic elements to perform operations to perform computational tasks. As an example, an FPGA can include several arrays of gates that can be configured to perform certain computational tasks in parallel. Thus, compared to software-driven host components, acceleration components can perform some operations in less time. In the context of the present disclosure, “acceleration” reflects its potential to accelerate functionality performed by host components.

[0031] Figure 1 A diagram illustrating an acceleration component 100 including a time-synchronized transport layer (TSL) component 130 is shown, according to one example. Components included in the acceleration component 100 can be implemented on hardware resources (e.g., logic blocks and programmable interconnects) of the acceleration component 100. The acceleration component 100 can include application logic 110 and a shell 120. The application domain hosts the application logic 110, which performs service-specific tasks (such as part of the functionality for ranking documents, encrypting data, compressing data, facilitating computer vision, facilitating speech conversion, machine learning, etc.). The shell 120 can be associated with resources corresponding to lower-level interface-related components that can generally remain the same in many different application scenarios. The application logic 110 can also be conceptualized as including an application domain (e.g., a “role”). The application domain or role can represent a portion of functionality included in a combined service that is rolled out across many acceleration components. Roles at each of a group of acceleration components can be linked together to create a group that provides service acceleration for the application domain.

[0032] In operation, in this example, the application logic 110 can interact with the shell resources in a manner similar to how a software-implemented application interacts with its underlying operating system resources. From an application development perspective, the use of common shell resources makes it unnecessary for developers to recreate these common components for each service.

[0033] The application logic 110 can also be coupled with the TSL component 130 via a transmit (TX) FIFO 132 for transmitting data to the TSL component 130 and a receive (RX) FIFO 134 for receiving data from the TSL component 130. The application logic 110 can also exchange metadata with the TSL component 130. In this example, the metadata can include transmit (TX) metadata and receive (RX) metadata. The TX metadata can include a destination ID, which can identify a particular acceleration component and a payload size. The RX metadata can include a source ID, which can identify a particular acceleration component and a payload size. Additional details regarding the metadata and operations of the TSL component 130 are provided later.

[0034] The enclosure resources included in the enclosure 120 can include a TOR interface 140 for coupling the acceleration component 100, a network interface controller (NIC) interface 150, a host interface 160, a memory interface 170, and a clock 180. The TOR interface 140 can also be coupled to the TSL component 130. The NIC interface 150 can be coupled to the application logic 110. A data path can allow traffic from the NIC or the TOR to flow into the acceleration component 100 and allow traffic from the acceleration component 100 to flow out to the NIC or the TOR.

[0035] The host interface 160 can provide functionality that enables the acceleration component 100 to interact with a local host component (not shown). In one implementation, the host interface 160 can use Peripheral Component Interconnect Express (PCIe) in conjunction with direct memory access (DMA) to exchange information with the local host component. The memory interface 170 can manage interactions between the acceleration component 100 and local memory, such as DRAM memory.

[0036] As noted earlier, the enclosure 120 can also include a clock 180. The clock 180 can include signal generators, including oscillators and phase-locked loops (PLLs), as needed. The clock 180 can also include hardware / software for allowing the clock 180 to update its time and thereby correct for any discrepancies with time managed by other clocks. The clock 180 can provide timing information to the TSL 130 and other components of the acceleration component 100. The enclosure 120 can also include various other features, such as status LEDs, error correction functionality, etc.

[0037] Multiple acceleration components like the acceleration component 100 can be configured to work in concert to accelerate a service. The acceleration components can use different network topologies to communicate with each other. Although Figure 1 A number of components of the acceleration component 100 are shown arranged in a certain way, but there can be more or fewer components arranged differently. In addition, the various components of the acceleration component 100 can also be implemented using other technologies.

[0038] Figure 2 A diagram of a TSL component 200 is shown, according to one example. The TSL component 200 can implement a time-synchronized transport layer protocol. The TSL component 200 can transmit and receive data via a TOR interface 208. In one example, the TOR interface 208 can be implemented as a 100 gigabit MAC interface. The TSL component 200 can include soft registers and configuration 210, a control logic state machine 220, a resilient buffer 230, a header queue 240, a data buffer 250, an Ethernet frame de-encapsulation 260, an Ethernet frame encapsulation 270, a TX metadata port 280, a RX metadata port 290, and a multiplexer 292. The soft registers and configuration 210 can include a connection table 212 and registers 214. The connection table 212 can include information about MAC and IP addresses of nodes (e.g., acceleration components or another type of endpoint). The connection table 212 can also include information about heartbeat data (described later) and any error related data.

[0039] Table 1 below shows an example data structure that can be used to establish a connection table.

[0040]

[0041] Table 1

[0042] As an example, if a network of acceleration components includes three acceleration components, each of the acceleration components can include a copy of the data structure shown in Table 1. The payload portion of the data structure can be used to characterize network latency from one node (e.g., an acceleration component) to another node. The control logic state machine 220 can be configured based on the payload. As explained later, a heartbeat corresponds to a message that is sent during a characterization phase and a standby phase of the protocol. In this example, the heartbeat messages can be used to measure time differences and pairwise latencies as needed to keep the nodes in the network synchronized with each other.

[0043] In one example implementation, the contents of the connection table (e.g., the connection table 212) can be the same regardless of whether the acceleration component is in a transmit phase or a receive phase. Additionally, although Table 1 shows certain contents of a data structure that can be used to establish a connection table, other approaches can also be used.

[0044] The registers 214 can be used to store configuration information about the acceleration component. Table 2 below shows one example of a data structure (in this example, in Verilog format) that can be used to generate configuration information for an acceleration component.

[0045]

[0046]

[0047] Table 2

[0048] In this example implementation, each node can have a unique connection table index (CTI). The value associated with the CTI can be used to read the connection table for a particular node. Each node can also include a CTI for the master node, which will be described in the context of the TSL protocol later.

[0049] In this example, using configuration parameters, the behavior of the TSL protocol can be customized for different network settings. As an example, by changing the number of heartbeat messages required to characterize the pairwise delay between nodes, the protocol can be customized for faster networks or slower networks. Additionally, efficiency improvements can be realized by configuring the use of jumbo Ethernet frames for data, such that nodes in the network can send and receive more than 1500 bytes / frame.

[0050] The control logic state machine 220 can be implemented in logic associated with the TSL component 200. As an example, the control and other logic associated with the control logic state machine 220 can be implemented as part of an FPGA that can be used to implement the acceleration component. The control logic state machine 220 can control the behavior of the endpoint / node (e.g., the acceleration component) in conjunction with the connection table 212 and the registers 214. As an example, the control logic state machine 220 can access the connection table stored in memory associated with the TSL component 200 to determine the IP address of the node to which data can be sent. As explained later, the TSL protocol can have phases, and the TSL component 200 can manage the transition from one phase to the next.

[0051] The control logic state machine 220 can be configured to handle various message types corresponding to the TSL protocol. As part of the TSL protocol, in one example, all messages can be encapsulated within an IPv4 / UDP frame. Table 3 below provides example message types and how they are enumerated in Verilog logic.

[0052]

[0053] Table 3

[0054] The control logic state machine 220 can also provide header decoding / verification for incoming messages (MSG) 222 and header generation for outgoing MSGs 224. Header decoding can involve decoding the message header and processing the information included in the header. The message header can be included as part of the UDP payload. Table 4 below provides an example structure for the message header.

[0055]

[0056] Table 4

[0057] The control logic state machine 220 can also be configured to track other aspects, including bookkeeping aspects. Table 5 below illustrates an example data structure for storing bookkeeping aspects of the TSL protocol.

[0058]

[0059] Table 5

[0060] As an example, as illustrated in Table 5 above, the bookkeeping aspects can involve tracking a timestamp and a sequence of data transmissions. For a sending node, the bookkeeping aspects can also include a megacycle ID, a sequence ID, and a number of bytes remaining to be sent. For a receiving node, the bookkeeping aspects can also include a megacycle ID, a sequence ID, a number of bytes still to be received, and an identity of the sending node. In other implementations, fewer or more bookkeeping aspects can be included.

[0061] Ethernet frames (e.g., jumbo Ethernet frames) can be provided to the TSL component 200 via the TOR interface 208. The Ethernet frames can be decapsulated via the Ethernet frame decapsulation 260. Table 6 below illustrates an example jumbo frame header that includes a TSL header (e.g., the TSL header illustrated in Table 4 above). The jumbo frame header can be decapsulated by the Ethernet frame decapsulation 260.

[0062]

[0063] Table 6

[0064] Decapsulation can result in the extraction of TSL message headers that can be buffered in the header queue 240. The header queue 240 can provide the headers to the control logic state machine 220, which can process them as previously described. The header queue 240 can also provide the headers to the elastic buffer 230. The functionality associated with the elastic buffer 230 and the multiplexer 292 is explained later.

[0065] Data extracted as a result of jumbo Ethernet frame decapsulation can be buffered in the data buffer 250, which can output the data to application logic associated with the acceleration component via the bus 252. Data received from the application logic via the bus 254 can be provided to the Ethernet frame encapsulation 270, which can encapsulate the data as jumbo Ethernet frames and provide them to the TOR interface 208.

[0066] Application logic (e.g., application logic 110 included as part of acceleration component 100) can also interact with TSL component 200 via TX metadata port 280 and RX metadata port 290. Application logic can send metadata information for transmission to application logic residing in another node via TX metadata port 280. Application logic can receive metadata information from application logic residing in another node via RX metadata port 290. As an example, metadata can include information about a transmission or reception schedule of a node participating in an acceleration service. As an example, the schedule can include information indicating to TSL component 200 who can send data to which node and the range of that data in the next megaword. If the node is a receiving node, the schedule can include the amount of data the receiving node should expect to receive in the next megaword and the identity of the sending node(s). Any other instructions or metadata can also be exchanged via the metadata ports. Although Figure 2 A particular number of components of TSL component 200 are shown arranged in a particular way, but there can be more or fewer components arranged differently.

[0067] Figure 3 A diagram of nodes in a network 300 for transmitting messages using a TSL protocol is shown according to one example. In this example, nodes 310, 320, and 330 can be coupled to a top-of-rack (TOR) switch 302, while nodes 340, 350, and 360 can be coupled to another TOR switch 304. Each node can include an acceleration component (A), a CPU (C), and a network interface controller (NIC). As an example, node 310 can include an acceleration component (A) 312, a CPU (C) 314, and a network interface controller (NIC) 316; node 320 can include an acceleration component (A) 322, a CPU (C) 324, and a network interface controller (NIC) 326; and node 330 can include an acceleration component (A) 332, a CPU (C) 334, and a network interface controller (NIC) 336. Similarly, in this example, node 340 can include an acceleration component (A) 342, a CPU (C) 344, and a network interface controller (NIC) 346; node 350 can include an acceleration component (A) 352, a CPU (C) 354, and a network interface controller (NIC) 356; and node 360 can include an acceleration component (A) 362, a CPU (C) 364, and a network interface controller (NIC) 366. Each acceleration component can correspond to an acceleration component 100 of Figure 1 and each of the acceleration components can include a TSL component 200 of Figure 2 Each node can include only FPGAs and ASICs and can not include any CPUs. Any arrangement of hardware computing components can be used as part of a node in Figure 3 .

[0068] TOR switch 302 can be coupled to a level one (LI) switch 306, and TOR switch 304 can be coupled to a LI switch 308. Level one switches 306 and 308 can be coupled to a level two switch 372. This is just one example arrangement. Other network topologies and structures can be used to couple nodes to communicate via the TSL protocol. In this example, IP routing can be used to transmit or receive messages between TOR switches. Each node or group of nodes can have a single "physical" IP address that can be provided by a network administrator. To distinguish IP packets that are destined for the CPU from packets that are destined for the acceleration components, UDP packets can be used that have a specific port used to designate the acceleration component as the destination. As noted earlier, the nodes can use the TSL protocol to communicate with each other.

[0069] The TSL protocol operates in stages 400 as shown in Figure 4 The possible transitions between stages are indicated by arrows. Example stages 400 include a startup stage 410, a characterization stage 420, a standby stage 430, a preparation stage 440, a data transfer stage 450, an error stage 460, and a reset stage 470. In this example, all nodes start in the startup stage 410. In this stage, each client (e.g., application logic) on a node can configure its TSL component (e.g., TSL component 200) before enabling it, in this example. The configuration process can include setting the MAC / IP addresses of all participating nodes, the UDP port they should use, the master node whose clock will be used as the global clock, and the like. In one example, the master node can be selected based on a clock standard associated with the nodes in the network during the characterization stage (explained later). As an example, the node whose clock has the smallest time drift can be selected as the master node. As explained earlier with respect to Figure 2 The application logic can use the TX metadata port and the RX metadata port to configure at least some of these aspects. Other aspects can be configured via connection tables and registers. In this example, the configuration should not change after the TSL protocol starts operating.

[0070] The characterization stage 420 can include measuring a delay associated with data transfers between the nodes in the network. As an example, node A can characterize at least one delay value by sending a message to node B, receiving the message back from node B, and measuring the time it took for the process. The delay value can be an average, minimum, or maximum. Continuing with the example from Figure 4 In the characterization stage 420, each node can periodically send a HEARTBEAT message to all other nodes. Upon receiving a HEARTBEAT message, a node responds with a HEARTBEAT-ACK message.Figure 5 An example of message exchange 500 between the primary node A, node B, and node C during the characterization phase 420 is shown. Upon receiving the HEARTBEAT-ACK message, a node calculates the one-way delay between itself and another node. In one example, assume that the local time of a node is denoted by t, while the global time (the time maintained by the master node (e.g., primary node A)) is denoted by T. Further, assume that a node (e.g., node B) sends a HEARBEAT / SYNC message to the master node (e.g., primary node A) at time t0, which is received by the master node at time T0. Further, assume that the master node responds by sending a HEARBEAT-ACK / SYNC-ACK message at time T1, and the node (e.g., node B) receives the HEARBEAT-ACK / SYNC-ACK message at time t1. Assuming an approximately constant symmetric delay, let At = T - t be the difference between the global time and the local time, d be the one-way delay between the non-master node and the master node, and ε be the variation in the delay, the one-way delay can be calculated by each node based on the following equation:

[0071]

[0072] The samples of the delay can be accumulated until the number of samples reaches a threshold value configured by the application logic (or, similarly, a client). The average and minimum delay values can be recorded by the TSL component (e.g., as part of the bookkeeping data structure described previously). In this example, this concludes the characterization of the link between two particular nodes. When a node completes the characterization of all links between itself and any other node, it takes the maximum of all average delays and the minimum of all minimum delays, and then sends a STANDBY message to the master node. Upon receiving the STANDBY message, the master node can update the maximum of all average delays and the minimum of all minimum delays of its own, and then it can respond with a STANDBY-ACK message. Upon receiving the STANDBY-ACK message, the non-master node can transition to the standby phase 430. After receiving the STANDBY message from all non-master nodes, the master node can transition to the standby phase 430. The client (e.g., application logic) can reset the nodes during the characterization phase 420, and thereby transition the nodes to the reset phase 470. The reset nodes can eventually transition back to the start-up phase 410.

[0073] Still referring to Figure 4Once the TSL component associated with a node (e.g., the acceleration component) transitions to the standby phase 430, it is ready to transmit data. All nodes continue to send HEARTBEAT, HEARTBEAT-ACK, STANDBY, STANDBY-ACK messages in the same manner as they did during the characterization phase 420. Figure 6 An example message exchange 600 between the primary node A, node B, and node C during the standby phase 430 is shown. A client can initiate a data transfer on a node by sending metadata (e.g., via the TX metadata port 280). A node that receives such metadata can transition to the prepare phase 440. A client (e.g., application logic) can reset a node during the standby phase 430 and thereby transition the node to the reset phase 470. A reset node can eventually transition back to the start phase 410.

[0074] With continued reference to Figure 4 During the prepare phase 440, the primary node can send a PROPOSAL message to all non-primary nodes. Figure 7 An example message exchange 700 between the primary node A, node B, and node C during the prepare phase 440 is shown. The PROPOSAL message can contain two values. The first value can correspond to a global synchronization time at the start of the data transfer megacycle. In one example, this is a time in the future with enough time for all nodes to synchronize before it arrives. The other value can relate to the period of the megacycle. In this example, the period of the megacycle can be calculated based on the characterization of the distribution of delays in the network and a range of configurations that can be modified during the start phase. As an example, the period of the megacycle can be chosen to be tight enough to keep all links busy, but also loose enough to tolerate unexpected delays in the network. A timer can be enabled after the primary node has sent all the PROPOSAL messages.

[0075] Upon receiving the PROPOSAL message, in this example, a node (e.g., node B) attempts to synchronize its local clock with the primary node (e.g., primary node A) shown in Figure 7 by sending a SYNC message to the primary node. The primary node responds to the SYNC message using a SYNC-ACK message. In this example, upon receiving the SYNC-ACK message, the node calculates the delay between itself and the primary node and the local time difference. If the delay exceeds a configurable margin of the average delay, the attempt is discarded, and the node sends a new SYNC message to the primary node. Otherwise, the attempt is accepted. The node (e.g., node B) then updates its local clock and sends a TX-READY message to the primary node (e.g., primary node A) as shown in Figure 7

[0076] ​If the master node does not receive a TX-READY message from all non-master nodes when the timer expires, the proposal is discarded. The master node sends a new PROPOSAL message with a new TX-MC start time and period, and then resets the timer. When all non-master nodes receive the new PROPOSAL message, they resend the SYNC message and the TX-READY message. If the master node has received a TX-READY message from all non-master nodes before the timer expires, the proposal is accepted. As shown in Figure 7 The master node (e.g., primary node A) then sends a PROPOSAL-CONFIRM message to all non-master nodes to inform them that the proposal has been accepted. Thereafter, all nodes stop sending messages in order to drain the network of messages exchanged until the preparation phase 440. When all nodes reach the specified mega-cycle start time, they transition to the data transfer phase 450. Since all nodes have synchronized their local clocks with the master node, such transition occurs at almost the same real-world time.

[0077] Figure 8 An example of message exchange and data transfer 800 between primary node A, node B, and node C during the data transfer phase 450 is shown. As shown in Figure 8 The data transfer phase 450 is divided into data transfer mega-cycles (TX-MC), as shown in

[0078]

[0079]

[0080] Table 6

[0081] As Figure 8As shown in , in this example, after sending all DATA messages scheduled for a particular megacycle, the node waits until the local time reaches the end of the current megacycle (TX-MC) period, and then it starts the next megacycle. In this example, the megacycle keeps progressing until the client sends "end" metadata, or the node detects an error. In this example, the megacycle period only restricts the node to being a transmitter. Correspondingly, the receive megacycle period (RX-MC) constrains the node to being a receiver, but in this example, there is no explicit time constraint for the RX-MC. Whenever a node that is not in a receive megacycle period (RX-MC) receives a DATA message that meets the following criteria: the TX-MCID is greater than the previously completed RX-MCID, and the sequence ID is 0, a new receive megacycle period (RX-MC) begins.

[0082] After a new receive megacycle (RX-MC) begins, the node expects consecutive DATA messages from the same sender until the last DATA message arrives. When the last DATA message arrives, the current receive megacycle (RX-MC) ends and the node sends a DATA-ACK message to the sender node. The node also sends RX metadata to the client when the last DATA message is received. If a node receives discontinuous DATA messages, it transitions to the error phase 460 and stops. The timeout of the DATA-ACK message causes all nodes to transition to the error phase 460. Although Figure 4 A certain number of phases are shown, but more or fewer phases may be included as part of the TSL protocol.

[0083] Figure 9 The effect of time drift 900 of a local clock relative to a master clock according to one example is shown. As shown in the figure above, since local time is counted separately at each node and the clock at each node may drift over time, synchronized nodes may gradually lose synchronization with the master node. In addition, non-TSL packets may exist in the network (e.g., broadcast Ethernet frames when the switch does not know the MAC address of certain nodes). Non-TSL packets interfere with the megacycle (TX-MC) time table and increase errors when measuring delay / time difference.

[0084] In order to keep the nodes in the network synchronized, and to tolerate changes introduced by non-TSL packets, periodic resynchronization can be added to the TSL protocol. Figure 10 As shown via message exchange 1000, at the end of each TX-MC, a non-master node (eg, Node B or Node C) may send a SYNC message to the master node (Primary Node A) to resynchronize. Figure 10The example in FIG. 10 shows that the clocks of nodes B and C have started to run faster than the global clock associated with the primary node A. In this example, it is assumed that the clock of node C is much faster than the global clock (managed by the primary node A), and that the clock of node B is a little faster than the global clock. In this example, during megacycle t+1, node C synchronizes its clock with the clock of the primary node A, while during megacycle t+2, node B synchronizes its clock with the clock of the primary node A. In one example, to avoid saturating the ingress link of the master node, non-master nodes only send SYNC messages when the TX-MCID is a multiple of its node ID.

[0085] While re-synchronization between nodes can improve the performance of the TSL protocol, there can be other aspects that can be improved. As an example, Figure 11 The slack is shown from the perspective of the messages 1100 sent and received by node B. Node B waits for an ACK message from the primary node A before starting the next megacycle. The time period between the last transmission from node B to the primary node A and the receipt of the ACK message from the primary node B shows the inefficiency of maintaining the slack. This is because, as part of the next megacycle, node B waits to receive the ACK message from the primary node A before starting the data transmission to node C.

[0086] Figure 12 One way to reduce the slack effect is shown in FIG. 12. Thus, Figure 12 The messages 1200 sent and received by node B in a way to reduce the slack are shown. In this case, node B does not wait to receive the ACK message from the primary node A, and starts the data transmission of the next megacycle before receiving the ACK message.

[0087] Figure 13 Another example of messages 1300 with slack is shown. This figure shows a different traffic pattern between the nodes. Thus, in this case, node C transmits messages to node B during the megacycles starting from time t. During the next megacycles, node C does not need to wait to receive the ACK message from node B before transmitting to the primary node A. However, node C still waits to ensure Figure 13 the slack shown in FIG. 10. Figure 14 The same example of messages 1400 with a smaller slack is shown. In one example, the smaller slack can be implemented using a flexible buffer (e.g., the flexible buffer 230 of FIG. 8). Figure 2 The flexible buffer can be configured to allow a small number of DATA messages to be buffered until the end of the current receive megacycle (RX-MC). As an example, Figure 15 An example of a message exchange 1500 including the use of a flexible buffer is shown. This figure shows the overlap between the messages received by the receiver, however these messages can be processed using a flexible buffer. As an example, the flexible buffer can be used to buffer the messages 1502 received by the receiver. The messages 1502 can be processed by the receiver in the order they are received, or they can be processed in the order they are buffered.Figure 15 As shown in FIG. 2, during a data transmission meacrop cycle that includes message transmission from transmitter 0 to receiver, there can be interference in the network link. This can result in the receiver having an overlap in the reception of messages despite the inter-mecacrop cycle (MC) margin. The use of the elastic buffer delays the processing of the received DATA messages, so it can result in reinterleaving in the subsequent receive meacrop cycle (RX-MC). However, given that network interference occurs infrequently, the impact should be gradually absorbed by the margin between data transmission meacrop cycles (TX-MC).

[0088] In one example, the elastic buffer can be implemented as the elastic buffer 230 previously shown in FIG. 2. In this example implementation, the elastic buffer 230 can store only the headers associated with messages that need to be queued due to an overlap between messages being received by the node. As shown in FIG. 2, the headers in the header queue 240 that do not need to be buffered are passed to the RX metadata port 290 via the multiplexer 292. The buffered headers are passed via the multiplexer 292 when the client (e.g., application logic) is ready to process the headers. The elastic buffering can also be implemented in other ways. As an example, instead of storing the headers, pointers or other data structures pointing to the headers can be stored in the elastic buffer. Figure 2 Figure 2 In one example, the elastic buffer can be implemented as the elastic buffer 230 previously shown in FIG. 2. In this example implementation, the elastic buffer 230 can store only the headers associated with messages that need to be queued due to an overlap between messages being received by the node. As shown in FIG. 2, the headers in the header queue 240 that do not need to be buffered are passed to the RX metadata port 290 via the multiplexer 292. The buffered headers are passed via the multiplexer 292 when the client (e.g., application logic) is ready to process the headers. The elastic buffering can also be implemented in other ways. As an example, instead of storing the headers, pointers or other data structures pointing to the headers can be stored in the elastic buffer.

[0089] Figure 16 A flowchart 1600 of a method for exchanging messages between nodes using a TSL protocol is shown according to one example. Step 1610 can include each of at least a first node, a second node, and a third node synchronizing respective clocks to a common clock. In this example, each of the first node, the second node, and the third node can correspond to the acceleration component 100 of FIG. 1. The nodes can communicate with each other using an arrangement such as that shown in FIG. 2. In this example, a TSL component (e.g., the TSL component 200) can perform this step during the preparation phase of the TSL protocol previously described. The common clock can correspond to a clock associated with one of the nodes in the network, which can be selected as a primary or master node. As previously described, the synchronization can be performed by a node (e.g., node B in FIG. 1) by attempting to synchronize its local clock to the master node (e.g., node A in FIG. 1) by sending a SYNC message to the master node. Figure 1 Figure 3 In one example, the elastic buffer can be implemented as the elastic buffer 230 previously shown in FIG. 2. In this example implementation, the elastic buffer 230 can store only the headers associated with messages that need to be queued due to an overlap between messages being received by the node. As shown in FIG. 2, the headers in the header queue 240 that do not need to be buffered are passed to the RX metadata port 290 via the multiplexer 292. The buffered headers are passed via the multiplexer 292 when the client (e.g., application logic) is ready to process the headers. The elastic buffering can also be implemented in other ways. As an example, instead of storing the headers, pointers or other data structures pointing to the headers can be stored in the elastic buffer. Figure 7 Figure 7 ​​​Synchronization is achieved synchronously as shown in the primary node A) in the middle. The primary node can respond to a SYNC message using a SYNC-ACK message. In this example, upon receiving a SYNC-ACK message, a node can calculate its own delay and local time difference from the primary node (as explained previously). The node (e.g., node B) can then update its local clock to synchronize with the clock associated with the primary node (e.g., primary node A). Synchronization between nodes can also be performed using other techniques. As an example, nodes can use the Precision Time Protocol (PTP) standard to synchronize their clocks to a common clock.

[0090] Step 1620 can include each of the at least first node, second node, and third node scheduling data transmissions in the network in a manner such that at a particular time with reference to a common clock, each of the at least first node, second node, and third node is scheduled to receive data from only one of the first node, second node, or third node. In one example, this step can include performing data transmission phase 450. As noted previously, data transmission phase 450 is partitioned into data transmission megacycles (TX-MC). In each megacycle, a client (e.g., application logic) can send metadata to its own TSL component to indicate which node it wants to send data to and how many bytes it wants to send. In one example, a time slot schedule can be provided for each node such that each sender will only access one particular receiver during a time slot and each receiver will receive data from a particular sender during a time slot. Thus, in this example, the sender-receiver pair cannot change during a time slot. The sender-receiver pair can change between time slots. All time slots can be run lockstep according to a given schedule, thereby causing all sender-receiver node pairs to run using the same time slot schedule. In one example, the shell 120 associated with each acceleration component can have a programmed schedule such that at runtime, the acceleration component can transmit or receive data according to the schedule. In one example, the soft registers and configuration 210 associated with the TSL component 200 can be used to store the schedule for the nodes. The schedule for the time slots need not be static. As an example, changes to the schedule can be made dynamically by the nodes by coordinating the schedule among them. Although Figure 16 A particular number of steps are shown listed in a particular order, but there can be fewer or more steps, and the steps can be performed in different orders.

[0091] Figure 17A flow diagram 1700 illustrating a method for exchanging messages between acceleration components using a TSL protocol according to one example is shown. Step 1710 can include each of at least a first acceleration component, a second acceleration component, and a third acceleration component synchronizing a respective clock with a common clock, the common clock being associated with an acceleration component selected from among the at least first acceleration component, the second acceleration component, and the third acceleration component. In this example, each of the first acceleration component, the second acceleration component, and the third acceleration component can correspond to a node 100 of Figure 1 . These acceleration components can communicate with each other using an arrangement such as that shown in Figure 3 . In this example, a TSL component (e.g., TSL component 200) can perform this step during a prepare phase of the TSL protocol previously described. The common clock can correspond to a clock associated with one of the acceleration components in the network, which can be selected as a primary node or master node. The synchronization can be achieved by an acceleration component (e.g., node B in Figure 7 ) attempting to synchronize its local clock with a master node (e.g., primary node A shown in Figure 7 ) by sending a SYNC message to the master node. The master node can respond to the SYNC message using a SYNC-ACK message. Upon receiving the SYNC-ACK message, in this example, the acceleration component can calculate a delay between itself and the master device and a local time difference (as previously explained). The acceleration component (e.g., node B) can then update its local clock to synchronize with the clock associated with the master node (e.g., primary node A). The synchronization between acceleration components can also be performed using other techniques. As an example, the acceleration components can synchronize their clocks with the common clock using a Precision Time Protocol (PTP) standard.

[0092] Step 1720 can include scheduling data transfers in the network in a manner such that at a particular time with reference to a common clock, each of at least the first, second, and third acceleration components is scheduled to receive data from only one of the first, second, or third acceleration components. In one example, this step can include performing the data transfer phase 450. As noted earlier, the data transfer phase 450 is partitioned into data transfer megacycles (TX-MCs). In each megacycle, a client (e.g., application logic) can send metadata to its own TSL component to indicate which acceleration component it wants to send data to, and how many bytes it wants to send. In one example, a time slot schedule can be provided for each acceleration component, such that each sender will access only one particular receiver during a time slot, and each receiver will receive data from a particular sender during a time slot. Thus, in this example, the sender-receiver pair cannot change during a time slot. The sender-receiver pair can change between time slots. All time slots can run in lockstep according to a given schedule, such that the sender-receiver pairs of all acceleration components run using the same time slot schedule. In one example, the shell 120 associated with each acceleration component can have a programmed schedule such that at runtime, the acceleration component can transmit or receive data according to the schedule. In one example, the soft registers and configuration 210 associated with the TSL component 200 can be used to store the schedule of the acceleration components. The schedule of time slots need not be static. As an example, changes to the schedule can be made dynamically by the acceleration components by coordinating the schedule among them. Although Figure 17 Although a specific number of steps are illustrated in a particular order, there can be fewer or more steps, and the steps can be performed in a different order.

[0093] In summary, the present disclosure relates to a method in a network comprising at least a first node, a second node, and a third node. The method can include synchronizing, by each of at least the first node, the second node, and the third node, a respective clock to a common clock. The method can further include scheduling, by each of at least the first node, the second node, and the third node, data transfers in the network in a manner such that at a particular time with reference to the common clock, each of at least the first node, the second node, and the third node is scheduled to receive data from only one of the first node, the second node, or the third node.

[0094] Each of at least the first node, the second node, and the third node can be configurable to provide service acceleration for at least one service. Each of at least the first node, the second node, and the third node can be configurable to communicate using a time-synchronized transport layer (TSL) protocol. The TSL protocol can include a plurality of phases, including a characterization phase, a standby phase, a preparation phase, and a data transfer phase.

[0095] The characterization phase can include at least one of (1) determining a first set of latency values associated with data transfer from the first node to the second node or the third node, (2) determining a second set of latency values associated with data transfer from the second node to the first node or the third node, and (3) determining a third set of latency values associated with data transfer from the third node to the first node or the second node. A node from among at least the first node, the second node, and the third node can be selected as a master node, wherein the common clock is associated with the master node, and wherein the master node is configured to transition from the characterization phase to the standby phase upon receiving a standby message from each of the nodes in the network other than the master node.

[0096] Each node can include application logic configurable to provide service acceleration for at least one service, and wherein each node is configured to transition from the standby phase to the preparation phase upon receiving a request from the respective application logic to initiate data transfer. The scheduling can include one of dynamic scheduling or static scheduling.

[0097] In another example, the present disclosure relates to a system including a network configured to interconnect a plurality of acceleration components. The system can also include a plurality of acceleration components configurable to accelerate at least one service, wherein each acceleration component of the plurality of acceleration components is configured to synchronize a respective clock with a common clock, the common clock being associated with an acceleration component selected from among the plurality of acceleration components, and wherein each acceleration component of the plurality of acceleration components is configured to transfer data in the network in a manner such that at a particular time with reference to the common clock, each acceleration component of the plurality of acceleration components is scheduled to receive data from only one of the plurality of acceleration components.

[0098] Each of at least the first acceleration component, the second acceleration component, and the third acceleration component can be configurable to provide service acceleration for at least one service. Each of at least the first acceleration component, the second acceleration component, and the third acceleration component can be configurable to communicate using a time-synchronized transport layer (TSL) protocol.

[0099] The TSL protocol can include a plurality of phases, including a characterization phase, a standby phase, a preparation phase, and a data transfer phase. The characterization phase can include determining a latency value associated with data transfer within the network. A selected acceleration component can be designated as a primary acceleration component, and the primary acceleration component can be configured to transition from the characterization phase to the standby phase upon receiving a standby message from each of the plurality of acceleration components in the network except the primary acceleration component.

[0100] The first acceleration component can be configurable to begin transmission of messages to the second acceleration component upon receiving an acknowledgement from the second acceleration component indicating completion of a data transfer cycle, or to begin transmission of messages to the second acceleration component prior to receiving the acknowledgement from the second acceleration component indicating completion of a data transfer cycle. Each of the plurality of acceleration components can include a resilient buffer configured to allow receipt of messages from two other acceleration components in a single data receive cycle.

[0101] In another example, the disclosure relates to a method in a network including at least a first acceleration component, a second acceleration component, and a third acceleration component. The method can include each of the at least first acceleration component, the second acceleration component, and the third acceleration component synchronizing a respective clock to a common clock associated with a selected acceleration component from among the at least first acceleration component, the second acceleration component, and the third acceleration component. The method can further include each of the at least first acceleration component, the second acceleration component, and the third acceleration component scheduling data transfer in the network in a manner such that at a particular time with reference to the common clock, each of the at least first acceleration component, the second acceleration component, and the third acceleration component is scheduled to receive data from only one of the first acceleration component, the second acceleration component, and the third acceleration component.

[0102] Each of the at least first acceleration component, the second acceleration component, and the third acceleration component can be configurable to provide service acceleration of at least one service. Each of the at least first acceleration component, the second acceleration component, and the third acceleration component can be configurable to communicate using a time-synchronized transport layer (TSL) protocol. The TSL protocol can include a plurality of phases, including a characterization phase, a standby phase, a preparation phase, and a data transfer phase.

[0103] It should be understood that the systems, methods, modules, and components described herein are merely exemplary. Alternatively or additionally, the functions described herein can be performed, at least in part, by one or more hardware logic components. For example and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc. In an abstract, but still tangible sense, any arrangement of components that performs the same functionality is effectively "associated" such that the desired functionality is achieved. Hence, any two components herein combined to achieve a particular functionality can be seen as "associated with" each other such that the desired functionality is achieved, irrespective of architectures or intermediate components. Likewise, any two components so associated can also be viewed as being "operably connected," or "coupled," to each other to achieve the desired functionality.

[0104] The functionality associated with some of the examples described in this disclosure can also include instructions stored in non-transitory media. The term "non-transitory media" as used herein refers to any media that stores data and / or instructions that cause a machine to operate in a specific manner. Example non-transitory media include non-volatile media and / or volatile media. Non-volatile media includes, for example, a hard disk, a solid-state drive, a magnetic disk or tape, an optical disk or tape, flash memory, EPROM, NVRAM, PRAM, or other such media, or a networked version of such media. Volatile media includes, for example, dynamic memory, such as DRAM, SRAM, cache, or other such media. Non-transitory media is distinct from transmission media, although it can be used in conjunction with transmission media. Transmission media is used to transmit data and / or instructions to and / or from a machine. Example transmission media can include coaxial cables, optical cables, copper wires, and wireless media, such as radio waves.

[0105] Furthermore, those skilled in the art will recognize that boundaries between the functionality of the above described operations are merely illustrative. The functionality of multiple operations can be combined into a single operation, and / or the functionality of a single operation can be distributed in additional operations. Moreover, alternative embodiments can include multiple instances of a particular operation, and the order of operations can be altered in various other embodiments.

[0106] Although the present disclosure provides specific examples, various modifications and changes can be made therein without departing from the scope of the present disclosure as set forth in the following claims. Accordingly, the specification and drawings are to be regarded in an illustrative manner and not a restrictive one, and all such modifications are intended to be included within the scope of the present disclosure. Any benefits, advantages, or solutions to problems identified herein are not to be construed as critical, essential, or mandatory to any or all claims.

[0107] Furthermore, the term "a" or "an" as used herein is defined as one or more than one. Also, the use of introductory phrases such as "at least one" and "one or more" in the claims should not be construed to imply that the introduction of another claim element by the indefinite articles "a" or "an" limits any of the claims to applications embodying only one such element, even when the same claim includes the introductory phrases "one or more" or "at least one" and indefinite articles such as "a" or "an." This same applies to the use of "comprising" (including other variations such as "comprise" and "comprises"), "having" (including its variations), "including" (including its variations), and "containing" (including its variations) to describe the character of who elements are included in the claims.

[0108] Unless otherwise stated, terms such as "first" and "second" are used to arbitrarily distinguish one element from another. Thus, these terms are not necessarily intended to indicate temporal or other prioritization of such elements.

Claims

1. A method in a network comprising at least a first node, a second node, and a third node, the method comprising: synchronizing, by each of at least the first node, the second node, and the third node, a respective clock to a common clock; and scheduling, by each of at least the first node, the second node, and the third node, data transmission in the network in a manner such that at a particular time with reference to the common clock, each of at least the first node, the second node, and the third node is scheduled to receive data from only one of the first node, the second node, or the third node, wherein each of at least the first node, the second node, and the third node is configurable to communicate using a time-synchronized transport layer (TSL) protocol, the TSL protocol comprising a plurality of phases, the plurality of phases including a characterization phase, and the characterization phase including at least one of (1) determining a first set of delay values associated with data transmission from the first node to the second node or the third node, (2) determining a second set of delay values associated with data transmission from the second node to the first node or the third node, and (3) determining a third set of delay values associated with data transmission from the third node to the first node or the second node.

2. The method of claim 1, wherein each of at least the first node, the second node, and the third node is configurable to provide service acceleration to at least one service.

3. The method of claim 1, wherein the TSL protocol includes a standby phase.

4. The method of claim 3, wherein the TSL protocol includes a preparation phase and a data transmission phase.

5. The method of claim 1, wherein a node from among at least the first node, the second node, and the third node is selected as a master node, wherein the common clock is associated with the master node, and wherein the master node is configured to transition from the characterization phase to a standby phase upon receiving a standby message from each of the nodes in the network other than the master node.

6. The method of claim 4, wherein each node includes application logic configurable to provide service acceleration to at least one service, and wherein each node is configured to transition from the standby phase to the preparation phase upon receiving a request from the respective application logic to initiate data transmission.

7. The method of claim 1, wherein the scheduling includes one of a dynamic schedule or a static schedule.

8. A communication system comprising: a network configured to interconnect a plurality of acceleration components; and ​ The plurality of acceleration components are configurable to accelerate at least one service, wherein each acceleration component of the plurality of acceleration components is configured to synchronize a respective clock to a common clock, the common clock being associated with an acceleration component selected from among the plurality of acceleration components, and wherein each acceleration component of the plurality of acceleration components is configured to transmit data in the network in a manner such that at a particular time with reference to the common clock, each acceleration component of the plurality of acceleration components is scheduled to receive data from only one of the plurality of acceleration components, wherein each acceleration component of the plurality of acceleration components is configurable to communicate using a time-synchronized transport layer (TSL) protocol, the TSL protocol comprising a plurality of phases, the plurality of phases comprising a characterization phase and a standby phase, wherein the characterization phase comprises determining at least one latency value associated with data transmission within the network.

9. The system of claim 8, wherein each acceleration component of the plurality of acceleration components is configurable to provide service acceleration to at least one service.

10. The system of claim 8, wherein the TSL protocol further comprises a preparation phase and a data transmission phase.

11. The system of claim 10, wherein the selected said acceleration component is designated as a primary acceleration component, wherein the preparation phase comprises: The master acceleration component sends a global synchronization time to each acceleration component of the plurality of acceleration components in the network other than the master acceleration component, and wherein the global synchronization time indicates a future time at which data transmission is to begin.

12. The system of claim 10, wherein the selected acceleration component is designated as a master acceleration component, and wherein the master acceleration component is configured to transition from the characterization phase to the standby phase upon receiving a standby message from each acceleration component of the plurality of acceleration components in the network other than the master acceleration component.

13. The system of claim 8, wherein a first acceleration component is configurable to begin transmission of a message to a second acceleration component upon receiving an acknowledgement from the second acceleration component indicating completion of a data transmission cycle, or to begin transmission of a message to the second acceleration component prior to receiving an acknowledgement from the second acceleration component indicating completion of a data transmission cycle.

14. The system of claim 13, wherein each acceleration component of the plurality of acceleration components comprises a resilient buffer configured to allow receipt of messages from at least two other acceleration components during a single data reception cycle.

15. A method in a network comprising at least a first acceleration component, a second acceleration component, and a third acceleration component, the method comprising: synchronizing, by each acceleration component of at least the first acceleration component, the second acceleration component, and the third acceleration component, a respective clock to a common clock, the common clock being associated with an acceleration component selected from among the first acceleration component, the second acceleration component, and the third acceleration component; and Each of the first, second, and third acceleration components are configured to communicate using a time-synchronous transport layer (TSL) protocol, the TSL protocol including a plurality of phases, the plurality of phases including a characterization phase, the characterization phase including determining at least one latency value associated with data transmission within the network.

16. The method of claim 15, wherein each of the first, second, and third acceleration components are configured to provide service acceleration to at least one service.

17. The method of claim 15, wherein the TSL protocol includes a standby phase.

18. The method of claim 17, wherein the TSL protocol further includes a preparation phase and a data transfer phase.

Citation Information

Patent Citations

  • Position-based broadcast protocol and time slot schedule for a wireless mesh network

    EP2811796A1

  • Method of CATV cable same-frequency time division duplex data transmission

    US20110185394A1