Deploying shadow buffers to inter-node bump-on-the-wire in clock-synchronized edge-based networking functionality.

JP2026531647APending Publication Date: 2026-09-17CLOCKWORK SYSTEMS INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2026515751
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-11
Filing Date
2024-09-05
Publication Date
2026-09-17

Smart Images

  • Figure 2026531647000001_ABST
    Figure 2026531647000001_ABST
Patent Text Reader

Abstract

A Bump-on-the-Wire (BOTW) associated with the transmitting host receives data packets destined for the receiving host. The data packets are transmitted by the transmitting host, and the transmitting Bump-on-the-Wire is positioned on the data path between the transmitting and receiving hosts. The transmitting host, receiving host, transmitting Bump-on-the-Wire, and receiving Bump-on-the-Wire are clock-synchronized with each other. The transmitting BOTW records the transmission timestamp of the data packets. The transmitting BOTW receives the reception timestamp of the data packets, along with auxiliary information, from the receiving Bump-on-the-Wire associated with the receiving host. Based on the transmission timestamp, reception timestamp, and auxiliary information, the transmitting BOTW determines a congestion metric and sends a congestion signal to the transmitting host based on the congestion metric.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure generally relates to coordinated control of network transmission and network traffic within a data flow. [Background Art]

[0002] Modern Internet infrastructure typically comprises large-scale data centers that generate enormous volumes of network traffic. When demand is high, the output of a data center may be constrained (e.g., by the capacity of switches, gateways, etc.), and it may be necessary to throttle network traffic. Such temporary congestion scenarios can cause bottlenecks and lead to packet loss. To ensure successful packet transmission when faced with such situations, systems have been developed that transmit an acknowledgment from a receiving node to a transmitting node when a packet is received. However, these acknowledgments are inefficient in that they further contribute to additional network traffic. Furthermore, these acknowledgments are limited to functioning only in scenarios involving a single sender and a single receiver. Furthermore, if an acknowledgment is not received, the packet is simply retransmitted ad hoc, potentially flowing into the same congested switch and leading to the same dropping outcome, resulting in scenarios where the packet is delayed indefinitely or not even received by the destination. Additionally, these scenarios are rooted in congestion that has already occurred and are insufficient to prevent congestion from occurring in the first place.

[0003] Many sources and destinations of network traffic (e.g., data centers, cloud computing, etc.) are opaque, and countless hops occur inside the black box that transmits and receives network traffic. As a result, it becomes impossible to implement countermeasures for identifying the cause of congestion, or for mitigating or avoiding congestion. [Prior Art Documents] [Patent Documents]

[0004] [Patent Document 1] U.S. Patent No. 10,623,173 [Patent Document 2] U.S. Patent No. 11,632,225 [Overview of the project]

[0005] A system and method for deploying bump-on-the-wires to detect congestion and instruct congestion control are disclosed, even in opaque systems that cause network congestion.

[0006] In some embodiments, a bump-on-the-wire (BOTW) associated with a transmitting host receives data packets destined for a receiving host, the data packets are transmitted by the transmitting host, the transmitting bump-on-the-wire is positioned on the data path between the transmitting and receiving hosts, and the transmitting host, receiving host, transmitting bump-on-the-wire, and receiving bump-on-the-wire are clock-synchronized with each other. The transmitting BOTW records the transmission timestamp of the data packets. The transmitting BOTW receives the reception timestamp of the data packets, along with auxiliary information, from the receiving bump-on-the-wire associated with the receiving host. The transmitting BOTW determines a congestion metric based on the transmission timestamp, reception timestamp, and auxiliary information, and sends a congestion signal to the transmitting host based on the congestion metric.

[0007] In some embodiments, a bump-on-the-wire (BOTW) associated with a transmitting host receives data packets destined for a receiving host, the data packets are transmitted by the transmitting host, the transmitting bump-on-the-wire is positioned on the data path between the transmitting and receiving hosts, and the transmitting host, receiving host, transmitting bump-on-the-wire, and receiving bump-on-the-wire are clock-synchronized with each other. The transmitting BOTW generates a modified data packet by adding the transmission timestamp of the data packet to the data packet and transmits the modified data packet to the receiving bump-on-the-wire on the path to the receiving host. The receiving BOTW determines a congestion metric based on the transmission timestamp, receiving timestamp, and auxiliary information, and sends a congestion signal to the transmitting host based on the congestion metric. [Brief explanation of the drawing]

[0008] [Figure 1] Figure 1 shows an exemplary system environment for implementing a network camera and priority functions according to one embodiment of the present disclosure. [Figure 2] Figure 2 is a network traffic diagram illustrating an embodiment of the present disclosure in which multiple sending hosts transmit multiple data flows to a single receiving host. [Figure 3] Figure 3 is a network traffic diagram illustrating the timestamp operation on both the transmitting and receiving sides of data transmission according to one embodiment of the present disclosure. [Figure 4] Figure 4 is a data flow diagram showing network camera activity during normal operation and the location where anomalies are detected, according to one embodiment of the present disclosure. [Figure 5] Figure 5 is a network traffic diagram illustrating an embodiment of the present disclosure in which a receiving host receives both high-priority and low-priority traffic from a transmitting host. [Figure 6] Figure 6 is a data flow diagram showing a network camera activity for which priority is considered when determining network camera activity, according to one embodiment of the present disclosure. [Figure 7] Figure 7 is a flowchart illustrating an exemplary process for performing a Netcam activity according to one embodiment of the present disclosure. [Figure 8] Figure 8 is a flowchart illustrating an exemplary process for performing Netcam activities in multiple priority scenarios according to one embodiment of the present disclosure. [Figure 9] Figure 9 is a data flow diagram showing Netcam activity in which a shadow buffer is considered according to one embodiment of the present disclosure. [Figure 10] Figure 10 is a flowchart illustrating an exemplary process for performing Netcam activity in coordination with shadow buffer considerations according to one embodiment of the present disclosure. [Figure 11] Figure 11 is a data flow diagram illustrating an exemplary process for triggering congestion control activity using transmit bumps on the wire, according to one embodiment of the present disclosure. [Figure 12] Figure 12 is a data flow diagram illustrating an exemplary process for triggering congestion control activity using a receive bump on the wire, according to one embodiment of the present disclosure. [Figure 13] Figure 13 is a flowchart illustrating an exemplary process for generating congestion notifications by transmit bumps in a wire, according to one embodiment of the present disclosure. [Figure 14] Figure 14 is a flowchart illustrating an exemplary process for generating congestion notifications by receiving bumps in a wire, according to one embodiment of the present disclosure. [Modes for carrying out the invention]

[0009] The figures and the following description relate to preferred embodiments for illustrative purposes only. From the following description, alternative embodiments of the structures and methods disclosed herein should be readily recognizable as viable alternatives that can be adopted without departing from the principles of the claims.

[0010] A system and method for coordinating data flow control in the face of temporary congestion are disclosed herein. “Netcam” monitors network traffic between clock-synchronized transmitting and receiving hosts, which is part of the data flow. As used herein, “Netcam” is an abbreviation of “Network Camera,” and is a module that tracks network traffic and ensures that corrective action is taken if the traffic of the data flow within the clock-synchronized system is delayed beyond acceptable limits. Netcam instructs transmitting and receiving hosts to buffer copies of network traffic according to several parameters (e.g., buffering a certain number of packets, buffering packets for a rolling window of time). Buffers may be overwritten on a rolling basis when the parameters are met (e.g., overwriting the oldest packets when new packets are transmitted or received, and when the buffer is full). Netcam may cause all transmitting and receiving hosts to write buffer data where anomalies are detected, and may cause transmitting hosts to retransmit packets that have been written. Retransmissions may be subject to jitter (e.g., time delays between packet transmissions in a data flow), and if a transmission delay or failure occurs as a result due to a given packet transmission sequence, the jitter may nevertheless cause enough change to make a retransmission attempt successful. The Netcam may determine the need to write and retransmit packets in different ways depending on the priority of the data flow. The Netcam may instruct the receiving host's shadow buffer to monitor path utilization and capacity, and high utilization and / or low capacity may cause the Netcam to anticipate an impending anomaly and take corrective actions similar to those taken when the buffer is full.

[0011] Advantageously, the Netcam implementation disclosed herein enables both improved network transmission and forensic analysis. Improved network transmission arises from the ability to retransmit accurate packet sets from many machines without relying on acknowledgment packets that might be lost or missing across a complex web of machines, by attempting to buffer the most recent packet transmissions across all machines in the data flow. Furthermore, virtual machines may have bugs that are difficult to detect or isolate. Writing packet sequences associated with anomalies enables fault analysis, which can allow for the identification of faulty virtual machines. Additionally, using shadow buffers to predict anomalies can prevent scenarios where traffic becomes excessively congested, leaving some capacity on the path and allowing corrective actions to occur without interrupting traffic. Further advantages and improvements are evident from the following disclosures.

[0012] Figure 1 shows an exemplary system environment for implementing a networkcam and priority function according to one embodiment of the present disclosure. As shown in Figure 1, the networkcam environment 100 includes a transmitting host 110, a network 120, a receiving host 130, and a clock synchronization system 140. Only one of the transmitting host 110 and the receiving host 130 is shown, but this is simply for illustrative purposes and ease of illustration, and any number of transmitting and receiving hosts may be part of the networkcam environment 100.

[0013] The transmitting host 110 includes a buffer 111, a network interface card (NIC) 112, and a network camera module 113. The buffer 111 stores a copy of outbound data transmission until one or more criteria for overwriting or discarding packets from the buffer are satisfied. For example, the buffer may store data packets until the capacity is reached, at which point the oldest buffered data packets may be discarded or overwritten. Other criteria may include the lapse of time (e.g., discarding a packet after a predetermined time has elapsed from its transmission timestamp), the amount of buffered packets (e.g., starting to discard or overwrite the oldest packets as new packets are transmitted after a predetermined amount of packets have been buffered), and the like.

[0014] In one embodiment, the buffer 111 stores information related to a given outbound transmission rather than the entire packet. For example, a byte stamp may indicate an identifier of the packet and / or a flow identifier, and a timestamp at which the packet (or aggregate data flow) was transmitted, and may be stored in place of the packet itself. In such an embodiment, the stored information does not need to be overwritten, and may be stored in the persistent memory of the transmitting host 110 and / or the clock synchronization system 140. This embodiment is not mutually exclusive with respect to the buffer 111 storing copies of packets, and they may be employed in combination.

[0015] The NIC 112 may be any type of network interface card, such as a smart NIC. The NIC 112 interfaces the transmitting host 110 and the network 120.

[0016] The netcam module 113 monitors a data flow for a specific condition and triggers a function based on the monitored data. By way of example, in response to detecting network congestion, the netcam module 113 may instruct all hosts that are part of the data flow to perform one or more of various activities, including suspending transmission, capturing a snapshot of buffered data transmission (i.e., writing buffered data packets to persistent memory), and performing other coordinated activities. As used herein, the term data flow may refer to a collection of data transmissions between two or more hosts that are associated with each other. Further details of the netcam module 113 are described in greater detail below in connection with FIGS. 2-8. The netcam module 113 may be implemented in any component of the transmitting host 110. In one embodiment, the netcam module 113 may be implemented within the NIC 112. In another embodiment, the netcam module 113 may be implemented in the kernel of the transmitting host 110.

[0017] The network 120 may be any network, such as a wide area network, a local area network, the Internet, or any other conduit for data transmission between the transmitting host 110 and the receiving host 130. In some embodiments, the network 120 may be within a data center that accommodates both the transmitting host 110 and the receiving host 130. In other embodiments, the network 120 may facilitate cross-data center transmission over any distance. References to a data center are merely exemplary, and the transmitting host 110 and the receiving host 130 may be implemented on any medium including those that are not data centers.

[0018] The receiving host 130 includes a netcam buffer 131, a NIC 132, a netcam module 133, and a shadow buffer 134. The netcam buffer 131, NIC 132, and netcam module 133 operate similarly to the analog components described above with respect to the transmitting host 110. Buffer 131 may be the same size as or different from buffer 111, and additionally or alternatively, may store byte stamps of received packets. Any further distinctions between these components implemented in the receiving host compared to the transmitting host will be revealed based on the disclosures in Figures 2-8 below.

[0019] The shadow buffer 134 may be used to track data traffic in a way that enables early warning of when congestion is likely to occur. For example, when data traffic is buffered, congestion may occur when the buffer is full, and the congestion will prevent further data traffic from flowing until the congestion is cleared. The shadow buffer may increment its counter faster than the regular buffer (for example, it may increment by 1.1 when one unit of data is received in the regular buffer) and / or decrement its counter slower than the regular buffer (for example, it may decrement by 0.9 or 0.95 when one unit of data is cleared in the regular buffer). As used herein, the term regular buffer may refer to the activity of buffers 111 and / or buffer 131, and / or the activity of other buffers disclosed herein having similar functions to those of buffers 111 and / or buffer 131. Although only one shadow buffer 134 is shown in Figure 1, multiple shadow buffers may be employed at the receiving host, and each shadow buffer may be assigned to a different subset of data flows, such as individual data flows corresponding to the same application. Shadow buffers may be incremented / decremented at different rates (e.g., to indicate more congestion for lower-priority applications and less congestion for higher-priority applications). Alternatively, shadow buffers may be incremented / decremented at the same rate, but different thresholding may be applied to different applications regarding when a data flow should be considered to be facing congestion. Data buffered in the regular buffer contains data traffic received by the receiver (e.g., network packets), and the data is removed from the regular buffer once the data has been processed and / or routed to its next destination.The activities described herein of the Netcam module 113 and / or Netcam system 140 that operate with respect to the conditions that are met for the regular buffer may also be performed if the shadow buffer 134 is congested.

[0020] The Netcam system 140 includes a clock synchronization system 141. The Netcam system 140 can monitor data observed by Netcam modules implemented on the host, such as Netcam modules 131 and 133. The Netcam system 140 can detect conditions requiring action by the Netcam modules and can send commands to the affected Netcam modules to take coordinated action on a given data flow. The clock synchronization system 141 synchronizes one or more components of each host, such as the NIC, kernel, or any other component on which the Netcam modules operate. Details of clock synchronization are described in the shared U.S. Patent Document 1, published April 14, 2020, which is incorporated herein by reference in its entirety. Individual hosts are synchronized to the same reference clock with very high precision, enabling accurate timestamps between hosts regardless of host location, host bandwidth conditions, jitter, etc. Further details of the Netcam system 140 are disclosed below with reference to Figures 2–8. The Netcam system 140 is an optional component of the Netcam environment 100, and the Netcam modules of the transmitting and / or receiving hosts can operate independently of a centralized system, other than relying on a synchronized reference clock.

[0021] The Netcam environment 100 offers many advantages. The Netcam module is edge-based, considering that it can be run in the kernel or NIC (e.g., smart NIC) of a host (e.g., a physical host, a virtual machine, or any other form of host). In one embodiment, the Netcam functionality may run as an underlay, meaning it may run as a shim on a layer of the OSI system below the congestion control layer (e.g., layer 3 of the OSI system). The Netcam module and / or Netcam system 140 may instruct a host to perform actions when conditions are detected, such as pausing the transmission of data flow across affected hosts, taking snapshots (i.e., writing some or all of the buffered data, such as the last N bytes sent and / or bytes sent in the last S seconds, where N or S may be default values ​​or may be defined by the administrator), and any other actions disclosed herein (e.g., congestion signals are detected using shadow buffers). Further advantages and features are described below with respect to Figures 2-8.

[0022] Figure 2 is a network traffic diagram illustrating a scenario in one embodiment of the present disclosure in which multiple transmitting hosts transmit multiple data flows to a single receiving host. As shown in Figure 2, transmitting host 1 transmits data flow 211 to receiving host 200, transmitting host 220 transmits data flow 221 to receiving host 200, and any number of additional hosts represented by transmitting host 230 may transmit their respective data flows (represented by data flow 231) to receiving host 200. As shown in Figure 2, the individual data flows transmitted by each transmitting host are different, but this is merely for convenience, and two or more transmitting hosts may transmit data from the same data flow. Furthermore, a single transmitting host may transmit two or more different data flows to receiving host 200. Although only one receiving host is illustrated, a transmitting host may transmit data flows to any number of receiving hosts.

[0023] The operation of the Netcam module on the transmitting and receiving hosts will now be described with reference to Figure 3. Figure 3 is a network traffic diagram illustrating the timestamp operation on both the transmitting and receiving sides of data transmission according to one embodiment of the present disclosure. As shown in Figure 3, when the transmitting host 310 sends a packet to the receiving host 320, the Netcam module 113 on the receiving host 320 records a transmit timestamp 311. Similarly, when the receiving host 320 receives a packet, the Netcam module 133 on the receiving host 320 applies a receive timestamp 321. The timestamps reflect the time when the data packet was transmitted or received by the relevant components (e.g., NIC, kernel, etc.) on which the Netcam module is installed. The transmit timestamps may be stored in buffers 111 and 131, attached to the packet, transmitted for storage in the Netcam system 140, or any combination thereof.

[0024] Since the transmitting host 310 is synchronized to the same reference clock as the receiving host 320, the elapsed time between the time of the transmit timestamp 311 and the receive timestamp 321 reflects the one-way delay of a given packet. In one embodiment, when a given packet is received, the receiving host 320 sends an acknowledgment packet indicating the receive timestamp 321 to the transmitting host 310, and the netcam module 113 can calculate the one-way delay by subtracting the transmit timestamp 311 from the receive timestamp 321. Other means of calculating the one-way delay are within the scope of this disclosure. For example, the transmit timestamp 311 may be attached to the data transmission, and the receiving host 320 may thereby calculate the one-way delay without requiring an acknowledgment packet. In yet another example, the netcam modules of the transmitting and receiving hosts may send timestamps collectively or individually to the netcam system 140, from which the netcam system 140 can calculate the one-way delay. For convenience and brevity, the scenario in which the transmitting host 110 calculates the one-way delay based on the acknowledgment packet will be the focus of the following disclosure, but those skilled in the art will recognize that any of these calculation methods are equally applicable.

[0025] In one embodiment, the Netcam system then determines whether the one-way delay exceeds a threshold. For example, after calculating the one-way delay, the transmitting host 110 may compare the one-way delay to a threshold. The threshold may be predetermined or dynamically determined. A predetermined threshold may be set by default or set by an administrator. Different thresholds may be applied to different data flows depending on one or more attributes of the data flow, such as the priority of the data flow, as will be further described below. The threshold may be dynamically determined depending on any number of factors, such as dynamically increasing the threshold as congestion decreases or decreasing the threshold as congestion increases (for example, because congestion is likely to indicate a problem where it is not the cause of the delay or is a minor cause). In one embodiment, the threshold may depend on the distance between the transmitting host and the receiving host, and may therefore be set per host. In such an embodiment, the threshold may be a predetermined multiple of the minimum one-way delay between the transmitting host and the receiving host. That is, the minimum one-way delay would be the minimum time required for a packet to travel from the transmitting host to the receiving host. The multiple is typically 1.5 to 3 times the minimum value, but may be any multiple defined by the Netcam administrator. The threshold is equal to a multiple of the minimum one-way delay. In response to a determination that the one-way delay exceeds the threshold, the Netcam module 113 may instruct the transmitting host 110 to take one or more actions.

[0026] In additional or alternative embodiments, the decision of whether to take one or more actions may be performed using a separate measure of the state of the shadow buffer (e.g., shadow buffer 134). In short (further details are given below), during a given data flow, and in parallel with buffering data using the regular buffer, the netcam module 133 may instruct the shadow buffer 134 to be incremented for individual units of data traffic received by the receiving host 320. The netcam module 133 may define a dynamic drain rate, which is the rate at which the netcam module 133 instructs the shadow buffer 134 to decrement. The dynamic drain rate may be determined by the netcam module 133 based on the number of data removed from buffer 131 per unit time (e.g., multiplying by a factor that causes drains to occur later in the shadow buffer 134 than in buffer 131). The Netcam module 133 can calculate the dwell time as a function of the shadow buffer 134's counter and dynamic discharge rate (for example, the dwell time can be calculated by dividing the value of the shadow buffer's counter by the dynamic discharge rate). From this, the Netcam module 133 can determine that the shadow buffer's one-way delay is the actual one-way delay aggregated with the dwell time (determined from the transmit and receive timestamps described above). The shadow buffer's one-way delay can be compared to a threshold (in addition to or instead of the regular buffer's one-way delay) to determine whether to take one or more actions.

[0027] Whether driven by regular buffer or shadow buffer one-way delay, one or more of these actions may include pausing transmission from the transmitting host when the one-way delay is high, which reduces congestion and, in general, reduces packet loss on network 120. The pause may be a predetermined time or may be determined dynamically in proportion to the magnitude of the one-way delay. In one embodiment, the pause may be equal to the one-way delay or may be determined by applying an administrator-defined multiple to the one-way delay. In one embodiment, Netcam may determine if a previous pause was enforced and, if so, reduce the pause time based on the amount of previous pause time that has already elapsed since the previously acknowledged packet. Furthermore, a given data flow may not be the only data flow contributing to congestion, and therefore its pause time may be shorter than the one-way delay or the one-way delay threshold.

[0028] Another possible action is to write some or all of the buffered data packets (e.g., from either or both the sending and receiving hosts) to persistent memory in response to a one-way delay exceeding a threshold. Diagnostics can then be performed on the buffered data packets (e.g., to identify network problems). Further actions are described with respect to Figures 4-8.

[0029] In some embodiments, data flows may be associated with different priorities. The Netcam module may determine the priority of a data flow based on an explicit identifier (e.g., a traffic layer identifier in the data packet header) or based on inference (e.g., based on a heuristic in which rules are applied to the packet header and / or payload to determine the priority type). As used herein, priority refers to a priority scheme in which types of data packets should be allowed to be transmitted and should be suspended during periods of congestion. The priorities disclosed herein are considered in the context of selecting which packets to transmit during network congestion, thereby avoiding the need for link underutilization or explicit bandwidth allocation.

[0030] To prioritize high-priority packets, a higher one-way threshold may be assigned to high-priority traffic, and a threshold relatively lower than the higher one-way threshold may be assigned to low-priority traffic. These thresholds may be used to compare with either or both of the shadow buffer one-way delay and / or the regular buffer one-way delay. Thus, for the Netcam module to detect anomalies, low-priority packets need to detect lower one-way delays, while high-priority packets only need to violate a higher one-way delay threshold to be detected as anomalies, so low-priority packets are detected as anomalies more frequently than high-priority packets. Following the above description of determining the one-way threshold for a given host, different one-way thresholds may be applied, depending on the priority, to different data packets transmitted or received by the same host. In the priority embodiment, the one-way threshold may be determined in the manner described above (e.g., by applying a predetermined multiple to the threshold), and the determination is further affected by applying a priority multiple. The priority multiple may be set by the administrator for any given type of priority, but higher for higher priorities and lower for lower priorities. Priorities do not need to be binary; any number of priority hierarchies can be established, each corresponding to a different type or multiple types of data traffic, and each having a different multiplier. Priorities and their associated multipliers can change over time for a given data flow (for example, priority may be reduced if the data flow begins transmitting a different type of data packet that does not require high-latency transmission).

[0031] In addition to using priority multipliers for the one-way delay threshold and varying the one-way delay threshold based on the priority of a given packet or data flow in which the packet is transmitted, the Netcam module may handle the pause time of paused traffic during pause operation in a priority-dependent manner. Lower pause times may be assigned to higher-priority traffic, and relatively higher pause times may be assigned to lower-priority traffic, ensuring that lower-priority traffic is paused more frequently than higher-priority traffic during congestion periods, thereby ensuring that higher-priority traffic has more bandwidth available while lower-priority traffic is paused. Pause times may be determined in the same manner as above, but with an additional step of applying an additional pause multiplier to the pause time, using lower pause multipliers (e.g., multiples less than 1 such as 0.7x) for higher-priority traffic and higher pause multipliers (e.g., multiples greater than 1) for lower-priority traffic.

[0032] Priorities can be assigned in any number of ways. In one embodiment, one or more "carpool lanes" can be assigned that can be used by data flows with eligible priorities. For example, a "carpool lane" may be a bandwidth allocation that does not guarantee a minimum bandwidth for a given data communication but is accessible only by data flows that meet essential parameters. Exemplary parameters may include one or more priorities that are eligible to use the reserved bandwidth of a given "carpool lane". For example, a carpool lane may require that the data flow has at least a medium priority, and therefore both medium and high priorities are eligible in a three-priority system with low, medium, and high priorities. In another example, there may be multiple carpool lanes (e.g., a carpool lane accessible by both medium and high priority traffic, in addition to a carpool lane accessible only by high priority traffic).

[0033] In one embodiment, guaranteed bandwidth can be allocated to a given priority. For example, a high-priority data flow may be allocated a minimum bandwidth, such as 70 Mbps. In such an embodiment, any surplus unused bandwidth from the guaranteed bandwidth may be allocated to lower-priority data flows until the bandwidth is required by the data flow for which the bandwidth is guaranteed. The guaranteed bandwidth may be absolute or relative. A relative guarantee ensures that a given priority data flow receives at least a certain relative amount of bandwidth than a low-priority data flow. For example, a high-priority data flow may be guaranteed three times the bandwidth of a low-priority data flow, and a medium-priority data flow may be guaranteed twice the bandwidth of a low-priority data flow.

[0034] Returning to Figure 2, if two or more sending hosts send data from the same data flow, those nodes, in cooperation with any receiving hosts receiving data from the data flow, may be referred to as a “cluster.” In one embodiment, a data flow may be identified by a set of identifiers that, if all are detected, indicate that a data packet is part of a data flow. For example, the Netcam module of any host may determine a flow identifier that identifies the data flow to which a packet belongs based on a combination of source address, destination address, source port number, destination port number, and protocol port number. Other combinations of identifiers may be used to identify the data flow to which a packet is part. As mentioned above, all hosts in a cluster, regardless of their form (e.g., servers, virtual machines, smart NICs, etc.), are clock-synchronized to the same reference clock.

[0035] In a scenario where data flows 211 and 221 are the same data flow, the transmitting host 210, transmitting host 220, and receiving host 200 form a cluster. Following this example, buffering of data packets (across both regular and shadow buffers) may occur at a per-flow level across the host cluster. That is, one or more Netcam modules and / or Netcam systems 140 may record in the host buffers of the data flow all packets transmitted or received within any parameters that the buffers use to record and then overwrite data (e.g., recently transmitted packets, packets transmitted / received within a given time, etc.). Furthermore, a receiving node that receives packets of a data flow from multiple transmitting hosts (e.g., receiving host 200 that receives packets from transmitting hosts 210 and 220) may maintain a single shadow buffer for the data flow, or it may maintain one separate shadow buffer for each of transmitting host 210 and transmitting host 220. In one embodiment, an index of the time sequence relative to a reference clock is stored with the buffer data (for example, the transmit timestamp 311 and / or receive timestamp 321 are stored with the buffer data packets). Thus, transmitting hosts 210 and 220 can store data packets 111 sharing a given flow ID in a buffer, and receiving host 200 can store received packets in buffer 131. Alternatively or additionally, transmitted and / or received packets may be sent to a networkcam system 140, which can buffer the received data.

[0036] This advantage of buffering a specific amount of data on individual hosts in the cluster allows different functions of the host netcam module to be enabled in response to the detection of anomalies (e.g., the aforementioned conditions referred to with respect to Figure 2 above). Figure 4 is a data flow diagram illustrating netcam activity during normal operation and where anomalies are detected, according to one embodiment of the present disclosure. The data flow 400 reflects host activity and netcam activity (e.g., activities taken by the netcam module of the transmitting / receiving host or netcam system 140) during normal operation and “anomaly function” (i.e., actions taken when an anomaly is detected). The data flow 400 first shows the normal function in which a host transmits or receives a data flow (402), and the netcam module or system (generally referred to in this figure as “netcam”) determines whether an anomaly is detected (e.g., based on one-way delay, as described above) (404). If no anomalies are detected, assuming the buffers are full from previous storage of data packets, the hosts (e.g., in a cluster) will overwrite those buffers (406) (meaning, for example, overwriting the oldest packets as mentioned above, or following some other overwrite heuristic). Of course, if the buffers are not full, overwriting is unnecessary and the packets are stored in the free memory of the buffers. Normal functioning repeats unless an anomaly is detected.

[0037] An anomaly function occurs when an anomaly is detected. Different anomaly functions are disclosed herein, and data flow 400 focuses on illustrating a specific anomaly function that retransmits buffered data. In the case of sending / receiving data flow information by a host (e.g., in a cluster) (408), Netcam may detect an anomaly (410). As described above, an anomaly is detected when the one-way delay (e.g., delay of the shadow buffer and / or regular buffer) exceeds a threshold. In the case of a cluster, the threshold may differ between hosts in the cluster, depending on the distance between the sending and receiving hosts. In response to the detection of an anomaly, Netcam instructs all hosts in the cluster to store the buffered data (412). That is, if an anomaly occurs on even one host in the cluster, data from all nodes in the cluster is stored. This can occur by instructing hosts to store the buffered data (or its portion related to the data flow) in persistent memory, or by holding the buffered data in a buffer and pausing data transmission, or by combining them with different instructions for different hosts. If pausing is used, the pausing time may vary across different nodes in the cluster, as described above. Regardless of how the data is stored, the Netcam may jitter the retransmission timing (414). The time sequence of sending and receiving packets is reflected in the stored data packets. The Netcam may jitter the retransmission timing by altering the time sequence (e.g., creating a longer delay between previous time gaps between transmissions, or transmitting packets in a different order) (414). Jitter may occur according to a heuristic or it may be random. Jitter is applied when a previously attempted time sequence was the cause of failure (e.g., the previously attempted time sequence itself may cause too much transient congestion), and therefore jitter may result in a jitter-free retransmission failing in such a scenario. The Netcam then retransmits the buffered data (or part of it) (416).Furthermore, instead of isolating packets of data flows associated with a data flow or anomaly, it may be more convenient and computationally efficient to retransmit the entire buffer containing data unrelated to the data flow or anomaly. Normal functionality then resumes until another anomaly is detected.

[0038] Retransmission with jitter is just one example of an anomaly, and any number of anomalies can occur in response to anomaly detection. For example, in addition to or alternative to the anomaly shown in data flow 400, buffered data may be written to persistent memory and stored for forensic analysis. In such a scenario, the Netcam may, in response to anomaly detection, send an alert to the administrator and / or generate an event log indicating the anomaly. Other aforementioned anomalies are equally applicable. As an example of forensic analysis, a known type of attack against systems such as data centers is a timing attack. A timing attack may have a "signature" in that the intervals between packets of traffic can be learned (e.g., by using pattern recognition, etc., by training a machine learning model with timing patterns labeled by whether the timing pattern was a timing attack). Forensic analysis may be performed to determine whether the data was a timing attack. Timing attacks can be blocked (e.g., by the Netcam module 113 dropping data packets from the buffer if it determines that the buffered data represents a timing attack).

[0039] As mentioned above, buffered data may include byte stamps (in contrast to, or in addition to, buffered packets). Byte stamps can be used for anomaly analysis (e.g., forensic analysis, network debugging, security analysis). The advantage of using byte stamps over buffered data packets is that they save memory space and have lower computational processing costs. Byte stamps for the period corresponding to an anomaly can be analyzed to determine the cause of the anomaly. The trade-off when using byte stamps over buffered packets is that buffered packet data is more robust and may provide further insight into the anomaly.

[0040] Figure 5 is a network traffic diagram showing a receiving host receiving both high-priority and low-priority traffic from a transmitting host, according to one embodiment of the present disclosure. As shown in Figure 5, transmitting host 510 transmits high-priority data flow 511 to receiving host 500, and transmitting host 530 transmits low-priority data flow 531 to receiving host 500. If network congestion occurs and an anomaly is detected, the transmitting host can treat high-priority and low-priority traffic differently. In one embodiment, since the low-priority data flow 531 is associated with a lower one-way delay threshold than the high-priority data flow 511, transmitting host 530 detects network congestion earlier than transmitting host 510. Therefore, transmitting host 530 can take corrective action, such as pausing network transmission of the low-priority data flow 531, during the pause time, while transmission of the high-priority data flow 511 can continue because its higher one-way delay threshold has not yet been reached. If a high-priority data flow 511 reaches its high one-way delay threshold and a pause action is taken in response, the pause time may be shorter than that of a low-priority data flow 531. This ensures that the high-priority data flow 511 resumes sooner and under less congestion than the low-priority data flow 531 would face if it were not paused for an extra amount of time.

[0041] Similarly, with respect to the operation of shadow buffers, high-priority shadow buffers may be maintained separately by the receiving host 500 for high-priority data flows 511, and low-priority shadow buffers may be maintained separately by the receiving host 500 for low-priority data flows 531. Ejection rates may be weighted differently on a priority basis. For example, high-priority shadow buffers may have a higher ejection rate than the ejection rate used for low-priority shadow buffers, and therefore high-priority shadow buffers are less likely to cause anomaly detection than low-priority shadow buffers.

[0042] Although shown as two separate transmitting hosts, transmitting hosts 510 and 530 may be the same host that transmits both high-priority and low-priority traffic to receiving host 500. Thus, the same transmitting host may continue transmitting high-priority data flows 511 as usual while taking corrective action (e.g., suspending) in response to the detection of anomalies in low-priority data flows 531. A transmitting host may have multiple buffers 111, each buffer may correspond to a different priority of data.

[0043] Figure 6 is a data flow diagram showing netcam activity in which priority is considered when determining netcam activity, according to one embodiment of the present disclosure. Data flow 600 begins when one or more transmitting hosts (e.g., transmitting host 110) transmit the data flow and apply a transmit timestamp (e.g., transmit timestamp 311) (602). A receiving host (e.g., receiving host 130) receives the data flow and applies a receive timestamp (e.g., receive timestamp 321) (604). Then netcam activity occurs. As described above, netcam activity can occur in the transmitting host (e.g., by receiving an ACK packet indicating a receive timestamp and by using a netcam module to calculate the one-way delay), in the receiving host (e.g., when a transmit timestamp is included in the data flow and the netcam module calculates the one-way delay from it), in the netcam system 140, or in some combination thereof.

[0044] The Netcam determines the one-way delay of data packets in a data flow (606). As described above, the one-way delay calculation may depend on the priority of the data flow, and therefore different data flows may have different one-way delay thresholds ("priority thresholds"). The one-way delay may generally be determined from the packets and / or aggregated with the dwell time to form a shadow buffer one-way delay. The Netcam compares the determined one-way delay (or delay, if a shadow buffer one-way delay is used) to the respective priority threshold (608). In response to the determination that the one-way delay is greater than the threshold for a given priority data flow (610), an abnormal function is initiated. As shown in Figure 6, some abnormal functions may include one or more of the following: pausing the transmission of a data flow associated with a given priority (612), and / or storing a buffer data flow associated with a given priority (614) (e.g., for forensic analysis). As described above, the pause time may vary depending on the priority level of the data flow being paused.

[0045] Figure 7 is a flowchart illustrating an exemplary process for performing a Netcam activity according to one embodiment of the present disclosure. Process 700 may be performed by one or more processors (for example, based on computer-readable instructions for performing an action, stored in non-temporary computer-readable memory). For example, Netcam modules 113, 133, and / or Netcam system 140 may perform some or all of the instructions for performing process 700. Process 700 is described in relation to Netcam module 113 for convenience, but may be performed by any other Netcam module and / or system.

[0046] Process 700 begins with the transmitting host (e.g., transmitting host 110) recording a first default amount of transmitted network traffic of the data flow on a first cycle basis (e.g., recording to buffer 111) (702) and the receiving host (e.g., receiving host 130) recording a second default amount of received network traffic of the data flow on a second cycle basis (e.g., recording to buffer 131) (704), with the transmitting and receiving hosts being clock-synchronized (e.g., using a reference clock of clock synchronization system 141).

[0047] The Netcam module 113 monitors for anomalies in the data flow based on the timestamps of data packets in the network traffic (for example, by subtracting the transmit timestamp 311 from the receive timestamp 321 and comparing the result with a one-way delay threshold) (706). The Netcam module 113 determines whether an anomaly is detected during monitoring (for example, based on whether the comparison indicates that the one-way delay is greater than the threshold) (708). In response to the determination that no anomalies are detected during monitoring, the Netcam module 133 may passively allow the recorded transmitted network traffic and recorded received network traffic to be overwritten by newly transmitted network traffic and newly received network traffic, respectively (710) (for example, by overwriting the oldest recorded data packets with the most recent network traffic and continuing to repeat elements 702-708). In response to a decision that an anomaly has been detected during monitoring, the Netcam module 113 pauses the data flow and instructs the transmitting host to store the recorded transmitted network traffic in a first buffer and the receiving host to store the recorded received network traffic in a second buffer (712).

[0048] Figure 8 is a flowchart illustrating an exemplary process for performing Netcam activities in multiple priority scenarios according to one embodiment of the present disclosure. Process 800 may be performed by one or more processors (for example, based on computer-readable instructions for performing an action, stored in non-temporary computer-readable memory). For example, Netcam modules 113, 133, and / or Netcam system 140 may perform some or all of the instructions for performing process 800. Process 800 is described in relation to Netcam module 113 for convenience, but may be performed by any other Netcam module and / or system.

[0049] Process 800 begins with the Netcam module 113 identifying a first data flow between a first transmitting host (e.g., transmitting host 110) and a receiving host (e.g., receiving host 130), the first data flow having high priority (e.g., high-priority data flow 511), and the transmitting and receiving hosts being synchronized using a common reference clock (802). The Netcam module 113 (e.g., of a different transmitting host or the same transmitting host as transmitting host 110) identifies a second data flow (e.g., a low-priority data flow 531) between a second transmitting host and a receiving host, the second data flow having low priority (804), and the second transmitting host may be the same or a different host as the first transmitting host.

[0050] The Netcam module 113 assigns a first delay threshold to the first data flow based on its high priority and a second delay threshold to the second data flow based on its low priority (806), and the first delay threshold exceeds the second delay threshold. The Netcam module 113 monitors the first one-way delay of the data packets of the first data flow relative to the first delay threshold (808), and monitors the second one-way delay of the data packets of the second data flow relative to the second delay threshold (810). In response to the determination that the first one-way delay of the data packets of the first data flow exceeds the first delay threshold, the Netcam module 113 pauses the transmission of the data packets of the first data flow from the first transmitting host to the receiving host for a first time (812). In response to the determination that the second one-way delay of the data packets of the first data flow exceeds the second delay threshold, the netcam module 113 pauses the transmission of the data packets of the second data flow from the second transmitting host to the receiving host for a second time that exceeds the first time (814).

[0051] Figure 9 is a data flow diagram showing Netcam activity illustrating the consideration of shadow buffers according to one embodiment of the present disclosure. Data flow 900 begins with a transmitting host sending the data flow and applying a transmit timestamp (902), and a receiving host receiving the data flow and applying a receive timestamp (904). These activities are performed in the manner described above with respect to elements 602 and 604 in Figure 6. As mentioned with respect to Figure 1, in one embodiment, the receiving host maintains both one or more regular buffers and one or more shadow buffers, the regular buffers storing data packets as they are received, and the shadow buffers maintaining counters that increase as data packets are received and are emptied according to a dynamic emptiness rate (i.e., decremented according to a dynamic emptiness rate every unit time). Different shadow buffers may be used for different data flows on the same receiving host, and different data flows may have different priorities.

[0052] The shadow buffer can be idle or active. The netcam module 133 of the receiving host 130 may determine that the shadow buffer is active in response to receiving traffic for a data flow (i.e., the shadow buffer for that data flow transitions from idle to active). The netcam module 133 may determine that the shadow buffer is idle in response to determining that traffic is no longer being received. For example, traffic may be considered no longer being received for a data flow if at least a threshold time has elapsed since the last packet of the data flow was received. As another example, if traffic for a data flow is consistently received packet by packet over a unit of time, and then a unit of time has elapsed in which no packets are received for the data flow, the netcam module 133 may determine that traffic is no longer being received. Thus, the netcam module 133 may continue to switch the state of the shadow buffer for a data flow from idle to active and back again, depending on whether traffic is being received for the data flow. As will be further described below, the state of the shadow buffer is used by the netcam module 133 to determine other attributes related to the shadow buffer, such as the emission rate.

[0053] In 904, assuming the shadow buffer was idle, the netcam module 133 transitions the shadow buffer from idle to active in response to receiving the first packet of a data flow (905a) and increments a counter in the shadow buffer that indicates the unit of received data traffic (905b). If the shadow buffer is already active, 905a is not performed, but 905b continues as each unit of traffic (e.g., packets) is received. In one embodiment, the netcam module 133 increments the counter by multiplying the unit of received data traffic by a coefficient. For example, for all received packets, the counter may be incremented by multiplying the unit by a number greater than 1 (e.g., 1.01 or 1.1). In a specific example where multiple priorities exist, when a packet is received, the shadow buffer may be multiplied by 1.01 if the packet is from a high-priority flow and by 1.1 if it is from a low-priority flow. The higher the coefficient, the faster the shadow buffer counter will have a number of threshold exceedances that reflect anomalies (e.g., scenarios that warrant traffic interruption and / or corrective action).

[0054] The NetCam (i.e., the NetCam system 140, the NetCam module 133, or any of several distributed processes) performs the NetCam activities shown in the rightmost column of Figure 9. For convenience, the activities are referred to as being performed by the NetCam module 133, but distributed or centralized processing by the NetCam system 140 is equally possible.

[0055] The Netcam module 133 determines the one-way delay of data packets for each data flow (906) and the dynamic emission rate for each shadow buffer corresponding to each data flow (908). Although 906 and 908 are shown sequentially in Figure 9, they may be performed in parallel with each other or in the reverse order of those shown. Element 906 may occur at any point between the events shown in Figure 9 and the occurrence of 914. Element 906 may be performed in the same manner as described above with respect to 606 in Figure 6.

[0056] The Netcam module 133 may determine a dynamic evacuation rate based on the number of units of data removed from the regular buffer per unit time while the shadow buffer is active. That is, if 3 bytes are removed from the regular buffer for transmission to the next node in a data flow per microsecond, a rate of 3 per microsecond is the basis on which the dynamic evacuation rate is determined, and it is multiplied by a coefficient of less than 1 (e.g., 0.9 or 0.95) so that evacuations from the shadow buffer occur slower than evacuations from the regular buffer. The reason for decrementing the shadow buffer at a slower rate than the regular buffer is, again, to ensure that if an anomaly occurs in the regular buffer, it is detected first using the shadow buffer. The Netcam module 133 may select a coefficient to multiply the evacuation rate based on the priority of the data flow, where high-priority data flows have a higher evacuation rate (e.g., 0.95-0.99), and medium-priority and low-priority data flows have lower evacuation rates (e.g., 0.9-0.94 for medium-priority data flows and 0.85-0.89 for low-priority data flows).

[0057] The netcam module 133 may determine the dynamic emission rate at any frequency (cadence), such as each time a data packet is received by the receiving host 130, or at a slower frequency, such as each Nth data packet received in a given data flow. The netcam module 133 may limit the performance of determining the dynamic emission rate (908) to scenarios in which the shadow buffer is active. If the shadow buffer is idle, the netcam module 133 may use the last determined dynamic emission rate as the static emission rate used to decrement the shadow buffer until time comes for the shadow buffer to become active again, after which the netcam module 133 may recalculate a new dynamic emission rate.

[0058] The dynamic efflux rate is used by the Netcam module 133 for two purposes. Firstly, the dynamic efflux rate is used to decrement the shadow buffer counter over time. Secondly, the dynamic efflux rate is used to calculate the “dwell time.” As used herein, the term “dwell time” refers to a value that can be aggregated with the actual one-way delay of packets on a data flow as a congestion signal to determine if there is an anomaly in the data flow that requires corrective action.

[0059] The netcam module 133 determines the residence time as a function of the shadow buffer counter (e.g., a proxy for the length of the regular buffer, having some additional length based on incremental and emission multiplier factors) and the dynamic emission rate (910). In one embodiment, the netcam module 133 calculates the residence time by dividing the shadow buffer counter value by the dynamic emission rate.

[0060] The NetCam module 133 determines a congestion signal for data flows based on dwell time (912). In one embodiment, the NetCam module 133 determines the congestion signal by mathematically aggregating the one-way delay between the transmitting host and the receiving host along with the dwell time. Similar to calculating the dynamic emission rate and incrementing a counter, the NetCam module 133 may weight the dwell time by a coefficient. For example, the dwell time may be weighted according to the priority of the data flow, where a larger multiplier may be used for lower priority data flows and a smaller multiplier may be used for higher priority data flows (e.g., 1.01-1.05 for high priority data flows, 1.06-1.14 for medium priority data flows, and 1.15-1.30 for low priority data flows). This again results in higher priority data flows being affected less frequently than lower priority data flows, causing their congestion signals to reach the threshold that triggers corrective action sooner.

[0061] In a manner similar to that described in Figure 6 for elements 610-614, the netcam module 133 may determine (914) that a congestion signal exceeds a threshold (e.g., a priority-specific threshold similar to the threshold used for the regular buffer) and take corrective action. Corrective action may include storing data or a display of data for the associated data flow (916) and / or pausing transmission of the associated data flow (918).

[0062] Figure 10 is a flowchart illustrating an exemplary process for performing netcam activity in coordination with shadow buffer considerations according to one embodiment of the present disclosure. Process 1000 may be performed by one or more processors (for example, based on computer-readable instructions for performing an operation, stored in non-temporary computer-readable memory). For example, netcam modules 113, 133, and / or netcam system 140 may perform some or all of the instructions for performing process 1000. Process 1000 is described in relation to netcam module 133 for convenience, but may be performed by any other netcam module and / or system.

[0063] Process 1000 begins with the netcam module 133 maintaining a plurality of buffers at the receiving host (1002), the plurality of buffers including a regular buffer and a shadow buffer (e.g., buffer 131 and shadow buffer 134). In response to receiving a data flow from a transmitting host clock-synchronized with the receiving host using a common reference clock, the netcam module 133 stores a first representation of the data of the data flow in the regular buffer (e.g., a data packet or metadata corresponding to a data packet in buffer 131), transitions the shadow buffer from idle to active (e.g., if this is the start of traffic for a data flow since the last interruption of traffic), and increments a counter in the shadow buffer indicating a unit of received data traffic (e.g., a counter in shadow buffer 134 corresponding to the data flow) (1004).

[0064] The Netcam module 133 determines a dynamic emptiness rate based on the number of data units removed from the regular buffer per unit time while the shadow buffer is active (1006), and the shadow buffer returns to an idle state in response to an interruption in the receiving host receiving the data flow. The Netcam module 133 calculates the dwell time as a function of the shadow buffer counter and the dynamic emptiness rate (1008), and determines a congestion signal for the data flow based on the dwell time (e.g., a congestion signal used to detect anomalies in the same way as described with respect to 708 in Figure 7) (1010).

[0065] Figure 11 is a data flow diagram illustrating an exemplary process for triggering congestion control activity using transmit bump on the wire according to one embodiment. As shown in Figure 11, data flow 1100 illustrates the process of sending a congestion signal to a transmitting host 1110 using transmit BOTW, based on delay and other information associated with the receive bump on the wire (BOTW).

[0066] In some embodiments, a bump-on-the-wire (BOTW) can be used to trigger congestion control activity. A BOTW can be any device implemented between one or more transmitting and receiving hosts. Exemplary BOTWs include, but are not limited to, NICs, smart NICs, FPGAs, etc., and a BOTW can be a component, switching component, or other network component that performs any additional processing on the path from the transmitting host to the receiving host. Although only one transmitting and one receiving host are shown in Figure 11, any number of transmitting hosts can operate through a single BOTW, as long as all data on the path to the receiving host passes through the BOTW from each transmitting host. A BOTW is shown as an edge device near the host in question, but does not necessarily have to be an edge device and can be located anywhere on the path between the transmitting and receiving hosts.

[0067] BOTW can be used to perform congestion control on a transmitting host even in environments where the transmitting host cannot be directly controlled by a congestion control service. For example, an entity may deploy a server and want a third party to perform congestion control. Congestion control services may not be able to control the server and its components (e.g., network interface cards, queues, buffers, etc.), especially if the server handles highly sensitive processing. By implementing BOTW, which can disguise normal server interactions, BOTW can cause the server to perform congestion control activities by providing information that triggers such activities, rather than directly commanding the server to perform them. BOTW has a Netcam module 133 installed or is operablely connected (e.g., in conjunction with a Netcam system 140) to perform congestion control activities. In Figure 1, each element shown for the transmitting host 110 and receiving host 130 (e.g., buffers, Netcam module, shadow buffer, etc.) is installed within BOTW or is communicably connected to BOTW.

[0068] In some embodiments, BOTW cannot perform such activities because there is no backpressure mechanism for closed-off servers. In some embodiments, BOTW is lightweight and does not have sufficient buffering and queuing components, but in such lightweight scenarios, by applying a shadow buffer to BOTW and clock-synchronizing the host and BOTW using the aforementioned synchronization mechanism, BOTW can use one-way delay and shadow buffer to notify the server of congestion and allow the server to perform the necessary congestion control. The server is just one example, and other closed network components (e.g., switches with line cards that cannot access congestion control services) are also within the scope of this disclosure. With respect to host and server activities, BOTW can directly perform clock-synchronization activities, network control activities, networkcam activities, and other functions disclosed herein, as shown in Figure 1-10.

[0069] As shown in Figure 11, the transmitting host 1110 sends data packets to the receiving host 1150 using data flow 1100. The transmitting BOTW 1120 receives the data packets and records the transmission timestamp of the data packets. The transmitting host 1110 and / or transmitting BOTW 1120 may have all or a distributed function of the transmitting host 110 in Figure 1. The timestamps have a common reference point for the clocks of the transmitting host 1110, receiving BOTW 1140, receiving host 1150, and other components of the network 1130, as each component is clock-synchronized to a common reference clock based on the activity performed by the clock synchronization system 141. The data packets reach the receiving BOTW 1140 via the network 1130 and are further sent to the receiving host 1150. The receiving hosts 1150 and 1140 may have all or a distributed function of the receiving host 130 in Figure 1. The receiving BOTW 1140 obtains the reception timestamp and sends the reception timestamp to the transmitting host. The receiving BOTW 1140 may also transmit auxiliary information 1160 from the shadow buffer of the receiving BOTW 1140, which includes the current size of the shadow buffer and, optionally, the emission rate information of the receiving BOTW 1140 (this is optional because, for example, if the emission rate of the receiving BOTW is constant, the transmitting BOTW 1120 may already have the emission rate stored). Implementing a shadow buffer in the receiving BOTW has the advantage of being able to determine congestion signals before congestion actually occurs in the data flow, including the transmitting host 1110 (and other transmitting hosts that are part of the data flow) and the receiving host 1150.

[0070] The transmitting BOTW receives the received timestamp and auxiliary information from the receiving BOTW and calculates the congestion metric. The congestion metric may represent approximate congestion, average congestion, or other measures of congestion. The output may be a value between 0 and 1, for example, or a value within any continuous range. An exemplary formula for calculating the congestion metric (P) is as follows:

[0071]

number

[0072] Thereafter, P is the congestion metric, OWD is the one-way delay calculated in the manner described above based on the transmit and receive timestamps, SSB is the size of the receive BOTW shadow buffer, DR is the rate of the receive BOTW shadow buffer, OWDT is the one-way delay threshold (this threshold is described with respect to elements 608 and 610 in Figure 6, and element 914 in Figure 9), and MDT is the maximum delay threshold. In some embodiments, the user may specify a "rate limit or rate guarantee" for the flow at each receiving node. The flow rate is estimated by sampling a substream of packets (e.g., to avoid overestimating the rate due to bursts) and counting the number of bytes they carry within a given time. The Netcam then notifies the flow of congestion based on the difference between the current rate estimate and the user-defined "nominal rate" of the flow (e.g., MDT). Explicitly estimating the rate in this way improves the flexibility and versatility of the bandwidth partitioning function.

[0073] Another example of a congestion metric is:

[0074]

number

[0075] And,

[0076]

number

[0077] And,

[0078]

number

[0079] teeth,

[0080]

number

[0081] These are the minimum and maximum congestion metric values ​​that satisfy the condition.

[0082]

number

[0083] And,

[0084]

number

[0085] This can be configured by the user to suit various network conditions.

[0086] After calculating the congestion metric, the transmitting BOTW 1120 (or any Netcam module from any connected system) may transmit a congestion signal 1170 based on the congestion metric. For example, in response to determining that (OWD+SSB / DR-OWDT) is greater than MDT, the transmitting BOTW 1120 may transmit a congestion signal (such as the ECN shown in the figure, or other congestion signals) to the transmitting host 1110. In response to determining that (OWD+SSB / DR-OWDT) is not greater than MDT, the transmitting BOTW 1120 may refrain from transmitting a congestion signal. The congestion signal may simply indicate the name of the congestion metric, or it may indicate that congestion is occurring based on a decision made by the transmitting BOTW 1120 based on the congestion metric (for example, if the shadow buffer indicates congestion exceeding a threshold, as described above with respect to Figure 9). The congestion signal 1170 may be information that the transmitting BOTW 1120 has added to an acknowledgment packet sent to the transmitting host via the transmitting bump on the wire. Acknowledgment packets may be generated by the transmitting BOTW 1120, or they may be intercepted based on the reception of data packets along the path from receiving host 1150 to transmitting host 1110. Intercepted acknowledgment packets may be modified to indicate congestion (e.g., by changing an optional header value). Congestion signals may be explicit congestion notices (ECNs) (e.g., when transmitting host 1110 uses an ECN-responsive control algorithm such as DCQCN (Data Center Quantization Congestion Notice) or DCTCP (Data Center Transmission Control Protocol)).

[0087] Figure 12 is a data flow diagram illustrating an exemplary process for triggering congestion control activity using a receive bump on the wire according to one embodiment. As shown in Figure 12, data flow 1200 includes the transmission of a data packet from transmitting host 1210 to receiving host 1260. In many respects, data flow 1200 operates similarly to data flow 1100, and for convenience, details already described with respect to data flow 1100 that are identical in operation may be omitted. Transmit BOTW 1220 receives the data packet and obtains a transmit timestamp (denoted as TxTimeStamp). Transmit BOTW 1220 may append a transmit timestamp 1230 to the data packet (e.g., by using an optional header value in the packet). The data packet (e.g., modified to include the transmit timestamp 1230) may be transmitted to receiving BOTW 1250 on its path to receiving host 1260 via network 1240. The receiving BOTW1250 may calculate the congestion metric (for example, in a similar manner to the method described above for the transmitting BOTW1120). The receiving BOTW1250 may send a congestion signal to the transmitting host 1210 (for example, by modifying the acknowledgment packet sent by the receiving host 1260, or by sending its own acknowledgment packet, etc.).

[0088] In either scenario of Figure 11 or Figure 12 (or both), using BOTW can be advantageous, in particular, for establishing time boundaries in the network. Time boundaries are defined and described in detail in the sharing patent document 2, published on April 18, 2023, which is incorporated herein by reference in its entirety. By using BOTW, time boundaries can be established by a service (e.g., clock synchronization system 141), and different network controls can be applied to different boundaries.

[0089] Figure 13 is a flowchart illustrating an exemplary process for generating congestion notification by a transmit bump on a wire according to one embodiment of the present disclosure. Process 1300 may be performed by one or more processors that execute instructions causing a BOTW and a host to perform the operations described herein. Process 1300 begins with receiving a data packet destined for a receiving host by a transmit bump on the wire associated with a transmitting host, the data packet being transmitted by the transmitting host, the transmit bump on the wire being located on the data path between the transmitting host and the receiving host, and the transmitting host, receiving host, transmit bump on the wire, and receive bump on the wire being clock-synchronized with each other (e.g., using a clock synchronization system 141) (1310).

[0090] The transmit bump on the wire records the transmit timestamp of the data packet (1320) and receives the receive timestamp of the data packet along with auxiliary information from the receive bump on the wire associated with the receiving host (1330). The transmit bump on the wire determines the congestion metric based on the transmit timestamp, receive timestamp, and auxiliary information (1340) and transmits a congestion signal based on the congestion metric (1350).

[0091] Figure 14 is a flowchart illustrating an exemplary process for generating congestion notification by a receive bump in a wire according to one embodiment of the present disclosure. Process 1400 may be performed by one or more processors that execute instructions causing the BOTW and the host to perform the operations described herein. Process 1400 begins with receiving a data packet destined for a receiving host by a transmit bump on the wire, the data packet being transmitted by the transmit host, the transmit bump on the wire being deployed at a location on the data path between the transmit host and the receiving host, and the transmit host, the receiving host, the transmit bump on the wire, and the receive bump on the wire being clock-synchronized with each other (1410).

[0092] The transmit bump on the wire generates a modified data packet by adding the transmit timestamp of the data packet to the data packet (1420), and sends the modified data packet to the receive bump on the wire on the path to the receiving host (1430). The receive bump on the wire determines the congestion metric based on the transmit timestamp, receive timestamp, and auxiliary information (1440), and sends a congestion signal to the transmit host based on the congestion metric (1450).

Claims

1. The receiving of a data packet destined for a receiving host in a transmit bump on the wire associated with a transmitting host, wherein the data packet is transmitted by the transmitting host, the transmit bump on the wire is located on the data path between the transmitting host and the receiving host, and the transmitting host, the receiving host, the transmit bump on the wire, and the receive bump on the wire are clock-synchronized with each other. In the aforementioned transmit bump on the wire, the transmission timestamp of the data packet is recorded, The receiving bump on the wire associated with the receiving host receives the reception timestamp of the data packet along with auxiliary information. The transmission bump on the wire determines the congestion metric based on the transmission timestamp, the reception timestamp, and the auxiliary information, The transmission of a congestion signal from the transmit bump on the wire to the transmit host based on the congestion metric, A method implemented on a computer, including the following.

2. The aforementioned transmit bump on the wire is a smart network interface card (NIC). The method implemented in a computer according to claim 1.

3. The aforementioned transmit bump on the wire is a field-programmable gate array (GPGA). The method implemented in a computer according to claim 1.

4. The transmit bump on the wire and the receive bump on the wire define a time boundary different from other time boundaries defined by other bump on the wires implemented on the same network as the transmit bump on the wire and the receive bump on the wire. The method implemented in a computer according to claim 1.

5. The auxiliary information includes the size of the shadow buffer implemented on the receive bump on the wire, and the output rate of the shadow buffer. The method implemented in a computer according to claim 1.

6. The aforementioned congestion metric is, Determine the quotient obtained by dividing the size of the shadow buffer by the output rate of the shadow buffer. Determine the total by adding the one-way delay to the aforementioned quotient. Determine the difference obtained by subtracting the one-way delay threshold from the above total. The congestion metric is obtained by dividing the difference by the maximum delay threshold. It is calculated by, The method implemented in a computer according to claim 5.

7. The one-way delay of the data packet is calculated based on the transmission timestamp recorded by the transmit bump on the wire and the receive timestamp recorded by the receive bump on the wire and sent back. The computer-implemented method according to claim 6.

8. Transmitting the congestion signal includes adding the congestion signal to the acknowledgment packet sent to the transmitting host via the transmit bump on the wire. The method implemented in a computer according to claim 1.

9. The congestion signal is an explicit congestion notification. The method implemented in a computer according to claim 1.

10. The aforementioned congestion metric is the mean congestion. The method implemented in a computer according to claim 1.

11. A non-temporary computer-readable medium having instructions encoded in memory, wherein, when the instructions are executed, they cause one or more processors to perform an operation, and the instructions The receiving of a data packet destined for a receiving host in a transmit bump on the wire associated with a transmitting host, wherein the data packet is transmitted by the transmitting host, the transmit bump on the wire is located on the data path between the transmitting host and the receiving host, and the transmitting host, the receiving host, the transmit bump on the wire, and the receive bump on the wire are clock-synchronized with each other. In the aforementioned transmit bump on the wire, the transmission timestamp of the data packet is recorded, The receiving bump on the wire associated with the receiving host receives the reception timestamp of the data packet along with auxiliary information. The transmission bump on the wire determines the congestion metric based on the transmission timestamp, the reception timestamp, and the auxiliary information, The transmission of a congestion signal from the transmit bump on the wire to the transmit host based on the congestion metric, A non-temporary computer-readable medium containing instructions that command something.

12. The receiving of a data packet destined for a receiving host in a transmit bump on the wire associated with a transmitting host, wherein the data packet is transmitted by the transmitting host, the transmit bump on the wire is located on the data path between the transmitting host and the receiving host, and the transmitting host, the receiving host, the transmit bump on the wire, and the receive bump on the wire are clock-synchronized with each other. In the aforementioned transmit bump on the wire, a modified data packet is generated by adding the transmission timestamp of the data packet to the data packet, The modified data packet is transmitted to the receive bump on the wire along the path to the receiving host. The reception bump on the wire determines the congestion metric based on the transmission timestamp, reception timestamp, and auxiliary information, The receiving bump on the wire transmits a congestion signal to the transmitting host based on the congestion metric, A method implemented on a computer, including the following.

13. The aforementioned receiving bump on the wire is a smart network interface card (NIC). The computer-implemented method according to claim 12.

14. The aforementioned receiving bump on the wire is a field-programmable gate array (GPGA). The computer-implemented method according to claim 12.

15. The transmit bump on the wire and the receive bump on the wire define a time boundary different from other time boundaries defined by other bump on the wires implemented on the same network as the transmit bump on the wire and the receive bump on the wire. The computer-implemented method according to claim 12.

16. The auxiliary information includes the size of the shadow buffer implemented on the receive bump on the wire, and the output rate of the shadow buffer. The computer-implemented method according to claim 12.

17. The aforementioned congestion metric is, Determine the quotient obtained by dividing the size of the shadow buffer by the output rate of the shadow buffer. Determine the total by adding the one-way delay to the aforementioned quotient. Determine the difference obtained by subtracting the one-way delay threshold from the above total. The congestion metric is obtained by dividing the difference by the maximum delay threshold. It is calculated by, The computer-implemented method according to claim 16.

18. The one-way delay of the data packet is calculated based on the transmission timestamp attached to the data packet and the reception timestamp recorded by the reception bump on the wire. The computer-implemented method according to claim 17.

19. Transmitting the congestion signal includes adding the congestion signal to an acknowledgment packet sent to the transmitting host by the receiving bump on the wire, the receiving bump on the wire intercepting the acknowledgment packet and adding the congestion signal. The computer-implemented method according to claim 12.

20. A non-temporary computer-readable medium having instructions encoded in memory, wherein, when the instructions are executed, they cause one or more processors to perform an operation, and the instructions The receiving of a data packet destined for a receiving host in a transmit bump on the wire associated with a transmitting host, wherein the data packet is transmitted by the transmitting host, the transmit bump on the wire is located on the data path between the transmitting host and the receiving host, and the transmitting host, the receiving host, the transmit bump on the wire, and the receive bump on the wire are clock-synchronized with each other. In the aforementioned transmit bump on the wire, a modified data packet is generated by adding the transmission timestamp of the data packet to the data packet, The modified data packet is transmitted to the receive bump on the wire along the path to the receiving host. The reception bump on the wire determines the congestion metric based on the transmission timestamp, reception timestamp, and auxiliary information, The receiving bump on the wire transmits a congestion signal to the transmitting host based on the congestion metric, A non-temporary computer-readable medium containing instructions that command something.

Citation Information

Patent Citations

  • US10,623,173

  • US11,632,225