Systems, methods, and storage media for bandwidth estimation filtering based on packet loss patterns

By identifying and filtering specific packet loss patterns, and using a bulge detection algorithm and a loss bulge filter, the accuracy problem of bandwidth estimation in SD-WAN architecture is solved, enabling accurate bandwidth estimation and network traffic optimization in the presence of packet loss.

CN116996925BActive Publication Date: 2026-02-27HEWLETT PACKARD ENTERPRISE DEV LP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211305566.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-25
Filing Date
2022-10-24
Publication Date
2026-02-27
Estimated Expiration
2042-10-24

AI Technical Summary

Technical Problem

In SD-WAN architectures, bandwidth measurement is difficult to perform, especially when network paths cannot be directly measured. Existing methods such as PRM cannot accurately estimate bandwidth, especially in the presence of packet loss, leading to estimation errors.

Method used

A bandwidth estimation filtering technique based on packet loss patterns is adopted. By identifying specific packet loss patterns, a bump detection algorithm (BDA) is used to estimate bandwidth, and a loss bump filter (LBF) is used to ignore negligible packet loss, thus ensuring the accuracy of the estimation.

Benefits of technology

In situations where tail drop queues exist, it can accurately estimate the available bandwidth of network paths, support optimized routing and load balancing of network traffic, and improve the accuracy and efficiency of network traffic engineering.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116996925B_ABST
    Figure CN116996925B_ABST
Patent Text Reader

Abstract

Various embodiments of the present disclosure relate to bandwidth estimation filtering based on packet loss patterns. Systems and methods are provided for implementing filtering techniques that are capable of estimating available bandwidth (e.g., over a network path) in the presence of modest loss caused by certain queue management techniques. When there is packet loss over a network path due to certain types of packet queue transmission mechanisms, methods, or models, bandwidth estimation can be performed using a bump in the road algorithm (BDA). When the pattern of packet loss (according to its signature) is identified as a pattern that can be used to perform a BDA to accurately estimate available bandwidth over the network path, the BDA-based bandwidth estimation can be used to place / route and load balance network traffic, or to participate in network traffic engineering, take other network-related action(s), or be reported out. Otherwise, the bandwidth estimation is suppressed and not used.
Need to check novelty before this filing date? Find Prior Art

Description

BACKGROUND

[0001] Bandwidth measurement is an important part of any network traffic engineering solution, including solutions using software-defined wide-area networks (SD-WAN). Before deciding where to place / route and load balance network traffic, the framework needs to know how much bandwidth is available on each network path.

[0002] In closed systems, direct measurements can be collected for each of the network devices on the traffic path. However, in many cases, it is not possible to collect direct measurements. For example, network devices can be in different administrative domains, or can be hidden by tunnels or encapsulation. This is also the case for SD-WAN, where SD-WAN gateways try to direct traffic to the best path on the Internet. BRIEF DESCRIPTION OF DRAWINGS

[0003] The present disclosure is described in detail according to one or more various examples in connection with the accompanying drawings. These examples are described in connection with the pertinent art to provide a full contextual understanding of the concepts presented. The examples described are intended merely to be illustrative and not limiting.

[0004] Figure 1 A computing environment for implementing bandwidth estimation is illustrated in accordance with some examples of the present disclosure.

[0005] Figure 2 An example of one-way delay measurement according to a decreasing method bandwidth measurement use case scenario is illustrated, where no significant packet loss occurs on the network path being measured.

[0006] Figure 3 An example of bandwidth measurement according to a one-way delay use case scenario and loss associated with a tail-drop queue method is illustrated.

[0007] Figure 4 An example of a tail-drop filter defining loss caused by a tail-drop queue is illustrated.

[0008] Figure 5 An example of bandwidth measurement according to a one-way delay use case scenario and loss associated with a token bucket policer method is illustrated.

[0009] Figure 6 An example of a tail-drop queue resulting from active queue management is illustrated.

[0010] Figure 7 An overview of a process for determining bandwidth estimation in accordance with some examples of the present application is illustrated.

[0011] Figure 8 are example computing components that can be used in the various features described in the examples of the present disclosure.

[0012] Figure 9A block diagram of an example computer system in which various examples described herein can be implemented is depicted.

[0013] The accompanying drawings are not exhaustive and do not limit the present disclosure to the precise forms disclosed. DETAILED DESCRIPTION

[0014] As noted above, it is difficult to perform bandwidth measurements in an SD-WAN architecture. When it is not possible to measure directly, bandwidth estimation can be implemented from two endpoints, which can be controlled or otherwise used for measurement. For example, bandwidth estimation can be performed by probing a network path with specially crafted probe packets sent from one end of the network path to the other. The receiving endpoint measures the time of receipt of the specially crafted probe packets and any changes in the queue delay / time pattern to estimate path characteristics, such as path capacity, available bandwidth, or bulk transfer capacity.

[0015] In some examples, a probe rate model (PRM) (e.g., using research projects including PathChirp, PathCos++, SLDRT, etc.) can be used to estimate the available bandwidth of a network path. The PRM can create temporary congestion states (e.g., the number of data packets transmitted over the network exceeds a threshold to reduce the available bandwidth across the network, etc.). Once a temporary congestion state is created on the network path, the controller can measure the increase in the queuing delay of the probe packets to infer when the network path is in a congested state. By determining when the network path is in a congested state and when it is not, such methods can then calculate an estimate of the available bandwidth.

[0016] However, PRMs can have several limitations. For example, standard PRMs can use packet timing to infer congestion states and can assume that there is no packet loss on the network path. However, buffer areas on network paths are not infinite, so packet loss (lost / dropped packets) can occur if congestion is severe enough. In addition, some network devices react to congestion states by specifically creating packet loss and never increasing the queuing delay, which defeats the purpose of PRM methods and prevents them from measuring available bandwidth. Furthermore, some network links are unreliable and can randomly drop packets, which also hinders bandwidth estimation.

[0017] Examples of the present disclosure relate to a filtering technique that can estimate available bandwidth (e.g., available bandwidth over a network path) in the presence of modest losses caused by certain queue management techniques (e.g., tail-drop queues). As will be described in greater detail below, tail-drop queues are an example of a possible cause of dropped packets (packet loss). In some examples, a "signature" of packet loss caused by a particular queue management technique, such as a tail-drop queue, can be identified, and a bandwidth estimation can be performed. In this way, a bandwidth estimation can be performed in scenarios where classic / conventional PRM techniques fail (recall that PRM assumes no packet loss).

[0018] In particular, certain types of packet loss (or causes of packet loss) can be associated with certain patterns of packet loss. If a particular pattern of packet loss is identified, this identification of the type or class of packet loss can be a basis for determining whether to use an estimation made using a bump detection algorithm (BDA), which will be described in greater detail below. That is, when there is packet loss on a network path due to certain types of packet queue transmission mechanisms, methods, or models, a bandwidth estimation can be performed using a BDA, with the expectation that the estimated available bandwidth will be an accurate estimate. In other scenarios, a BDA-based bandwidth estimation can result in a false available bandwidth estimate (typically an overestimation of the available bandwidth). Thus, when a pattern of packet loss (according to its signature) is identified as a pattern that can be used to perform a BDA to accurately estimate the available bandwidth over a network path, a BDA-based bandwidth estimation can be used for placing / routing and load balancing network traffic, or used to participate in network traffic engineering, take other network-related actions, or reported out.

[0019] As will be described in greater detail below, a loss bump filter (LBF) can be used, where the LBF can refer to defining a region in a received probe or chirp sequence in which packet loss can be ignored. This region can be defined according to a BDA-selected packet, or according to a one-way delay (OWD) of a packet. However, losses can still occur at or beyond such packet / measurement points, and in such cases, the LBF can be used to measure these losses to determine whether there is a need to prevent use of a potentially false bandwidth estimate. In some examples, when such losses are indicative of negligible or acceptable packet loss relative to a threshold, e.g., below a threshold, the BDA bandwidth estimate can be deemed accurate, and can be reported to a user for subsequent use in routing / balancing traffic, etc. If not, the BDA-based bandwidth estimate is deemed inaccurate, and the BDA-based bandwidth estimate is not relied upon for routing / balancing traffic.

[0020] Figure 1A computing environment for implementing various bandwidth estimation processes is illustrated. As shown Figure 1 The sender computing device 110 transmits one or more data packets to the receiver computing device 120 over the network 130, as shown. The sender computing device 110 and the receiver computing device 120 (e.g., clients, servers, etc.) can be endpoints in a communication network that connect endpoints. Various sender computing devices can send data packets to various receiver computing devices, although a single sender computing device and a single receiver computing device are provided for simplicity of explanation.

[0021] Available bandwidth for sending network communications between the sender computing device 110 and the receiver computing device 120, e.g., a network path, can be determined. Network bandwidth can correspond to how much bandwidth is not being used and how much network traffic can be added to the network path. The network path can be composed of multiple wired or wireless communication links to transmit data over the network 130 in a given amount of time. Available bandwidth can be reduced by traffic that is already using the network path.

[0022] Network traffic engineering is one example of available bandwidth usage. A communication path between the sender computing device 110 and the receiver computing device 120 via the network 130 can include multiple possible paths from one endpoint to the other. Having multiple paths between endpoints can generally improve failure recovery capabilities and can increase network bandwidth.

[0023] In network traffic engineering, characteristics of network traffic and related network elements and their connectivity can be determined to help both design the network and direct traffic onto different paths in the network. In some examples, a primary path between the sender computing device 110 and the receiver computing device 120 can add a secondary path to use in the event of a failure of the primary path.

[0024] Network traffic engineering can consist of three parts. The first part is measurement, in which some properties of the traffic and / or network are measured. As described herein, the measurement part of network traffic engineering can incorporate available bandwidth determination(s) for various network paths. The second part is optimization, in which the optimal distribution of traffic is computed. The third part is control, in which the network is reconfigured to achieve the desired traffic distribution.

[0025] The network bandwidth estimation process includes another context for software defined networking (SDN). Software defined networking (SDN) can manage a network by defining an application programming interface (API). The API can allow a system to separate the data path (e.g., packet forwarding) and control plane (e.g., protocol intelligence) of a network element. In other words, a network controller, an entity outside of a network element, can have fine-grained control and visibility into that network element. This can be used by the network controller to dynamically change the policy of the network element or centralize the control plane and decision making of the network.

[0026] The SDN approach can be combined with network traffic engineering. For example, SDN APIs typically define both measurement and control, which can enable a network controller to measure network bandwidth or capacity and dictate the distribution of traffic through network traffic engineering.

[0027] One of the limitations of SDN is that it assumes tight coupling between the network controller and the network elements. This can work in small to medium scale communication networks, but generally does not scale to larger networks. If the network between the network controller and the network elements has limited performance (e.g., low bandwidth or high latency), the efficiency of the SDN process is reduced. Furthermore, the SDN approach generally does not allow for crossing management domain boundaries, as different entities can only trust controlled and limited interaction between each other.

[0028] Another example where available bandwidth can be used is in a software defined wide area network (SD-WAN). The SD-WAN process can be implemented in a computing environment that is distributed across multiple physical locations, including cases where the sender computing device 110 and the recipient computing device 120 are in different physical locations. Thus, the computing environment can include a set of local networks (LANs) that support the local physical locations of the entities and a set of wide area network (WAN) links that interconnect these local networks.

[0029] For example, a large bank or retailer can have multiple physical locations and branches. Each location or branch has a set of LANs. In a traditional configuration, all branches and locations are connected to a few central locations using dedicated WAN links (e.g., using routing technologies such as multiprotocol label switching (MPLS)), and the few central locations can be connected to an external network (e.g., the Internet, etc.) using one or more WAN links. The dedicated WAN links can be provided by a telecommunications company, which generally corresponds to high availability and quality of service guarantees, but also has a high monetary cost.

[0030] SD-WAN flow proposes to use SDN principles to manage WAN connections. This can provide centralized visibility and control of the WAN connections of an entity and reduce the cost of the WAN connections. SD-WAN can replace dedicated WAN links with tunnels over the Internet to reduce costs. In this case, each branch and location has a WAN link that connects to the Internet, typically using inexpensive consumer WAN technology (e.g., digital subscriber line (DSL) modem, cable modem, wireless 3G, etc.). The network can implement a special SD-WAN gateway at each branch and location to create a dedicated tunnel (e.g., virtual private network (VPN), etc.) to securely connect to other branches and locations through the WAN link and the Internet.

[0031] Another example where available bandwidth can be used is tunnel switching. When an SD-WAN gateway detects that a WAN link is down, the gateway can direct traffic from that WAN link to a tunnel that does not use that particular WAN link, also known as tunnel switching. The SD-WAN gateway can create parallel tunnels on the network 130 using each WAN link and then use network traffic engineering to direct traffic to the most suitable tunnel with the goal of optimally using the available network capacity. In some examples, the SD-WAN gateway can monitor the performance of each tunnel in terms of latency and throughput and then load balance traffic or map each traffic type to the tunnel that is most suitable for that traffic.

[0032] One component of such traffic engineering is a method to measure the performance of each tunnel. Each tunnel defines a network path across the network, and the tunneled packets are processed by multiple network elements. The network path used by the tunnel (e.g., outside the tunnel) and the network path inside the tunnel are logically different because they have different addresses. However, in some cases, the two network paths can go through the same network elements, have nearly identical performance, and their performance characteristics are strongly correlated. Thus, it is possible to measure the performance of a tunnel by measuring the network path outside the tunnel or inside the tunnel.

[0033] In some examples, it is difficult to determine bandwidth estimates for the tunnel switching process. For example, direct measurement or SDN methods cannot be used for these network paths because the vast majority of the network elements of the Internet path are in different administrative domains (e.g., various ISPs on the path), and it is difficult to obtain a complete list of all these elements (which is dynamic in most cases) and administrative access to them. In some examples, path measurements can be done via end-to-end network path estimation methods and / or by sending probe packets from one endpoint to another, from the sender computing device 110 to the receiver computing device 120.

[0034] In particular, a sender computing device 110 transmits one or more data packets to a receiver computing device 120 via a network 130. The sender computing device 110 and the receiver computing device 120 (e.g., clients, servers, etc.) can be endpoints in a communication network that connect endpoints. Various sender computing devices can send data packets to various receiver computing devices, although for simplicity of explanation, a single sender computing device and a single receiver computing device are provided.

[0035] The sender computing device 110 can generate a probe train designed to measure the network path and send probe packets to the receiver computing device 120 via the network 130. The receiver computing device 120 can receive the probe packets and estimate the available bandwidth.

[0036] Various methods can be implemented to determine bandwidth estimates and network path estimates. For example, when direct measurements of network elements are not possible (e.g., when they are in different administrative domains, etc.), the next best procedure can be end-to-end bandwidth estimation.

[0037] End-to-end network path estimation can include active probing using data packets transmitted in the network 130. For example, a sender computing device 110 at one end of the network path sends special probe packets to a receiver computing device 120 at the other end of the network path. These packets can be used only for estimating bandwidth and can not carry actual data beyond what is needed for the network path estimation itself.

[0038] The estimation procedure can also include passive measurements, or by measuring the queuing delay experienced by existing data transmitted on the network path, or by modulating that data to have specific characteristics. Another variation is single-ended measurement, where the method is initiated by the sender computing device 110 of probe packets that are reflected back to the sender computing device 110.

[0039] Different methods can estimate different characteristics of the network path. Bandwidth estimation is one subset of network path estimation. Path capacity is the maximum traffic bandwidth that can be sent if the network path is free (i.e., no any competing traffic). Available bandwidth (ABW) is the rest / residual path capacity, i.e., the capacity that is not currently used by other traffic. Bulk transfer capacity (BTC) is the bandwidth that a Transmission Control Protocol (TCP) connection would get if placed on this network path. Latency is the one-way delay from sender to receiver, and Round-Trip Time (RTT) is the two-way delay.

[0040] With active probing, the sender computing device 110 sends a series of specially crafted probe packet patterns to the receiver computing device 120. The packet patterns can be defined by the estimation method and can be designed to trigger specific behavior from network elements on the network path. For example, in many cases, the packet pattern is a probe sequence (described in more detail below, also referred to as a chirp sequence), in which the packets and the spacing between packets are intended to span various bandwidths across the probe packet pattern. The receiver computing device 120 can measure the time of receipt of the packets and compute the OWD for each packet (i.e., the time it takes for a packet to travel from the sender device to the receiver device). The receiver computing device 120 can examine the changes in the packet pattern. The estimation method uses a simplified network model to convert these measurements into estimates of various network path characteristics.

[0041] For bandwidth estimation, two main categories can include a probe gap model (PGM) and the aforementioned PRM. For the PGM, it is assumed that two closely sent packets will see their gap between them increase proportionally to the load on the heaviest queue, due to the queuing delay on that queue. For the PRM, it is assumed that when packets are sent at a rate lower than the bottleneck bandwidth (where the performance of the network / network path is limited due to lack of available bandwidth to ensure that the packets can be received by a receiver, such as the receiver computing device 120, the performance of the network / network path is limited phenomenon or state), the traffic pattern will mostly remain unchanged, while when packets are sent at a rate greater than the bottleneck bandwidth, those packets will suffer additional queuing delay due to the congestion state of the network path.

[0042] In practice, the PGM and PRM attempt to infer network path congestion by attempting to estimate the variation in queuing delay experienced by packets at different network elements in the network path. Queuing delay affects the time it takes for a packet to traverse the network path. The PGM and PRM can compare the one-way delays of various probe packets to estimate the variation in queuing delay. For example, with the PGM, two packets can be sent at a known sending interval. It is assumed that the measured receiving interval is the sum of the sending interval and the difference in queuing delay between the packets.

[0043] Another illustrative network bandwidth estimation process includes packet one-way delay (OWD). In the PRM, this method can measure the delay of received packets to determine network path bandwidth. The measurement for each packet uses the OWD process. The OWD corresponds to the difference in time between the time a packet is sent by the sender computing device 110 (e.g., sending time) and the time the packet is received by the receiver computing device 120 via the network 130 (e.g., receiving time). Some methods compare the OWDs of multiple packets.

[0044] In some examples, the OWD of a packet can correspond to the propagation delay of the network path, the transmission time of the slowest link in the path, and the cumulative queuing delay of all network elements in the path. The formula to determine the OWD for each packet can include:

[0045] OWD(i) = pd + st(size) + sum(qd(e,i))

[0046] where:

[0047] pd -> total propagation delay

[0048] st(size) -> slowest transmission time for the size of this packet

[0049] qd(e,i) -> queuing delay at element e for packet i

[0050] In some examples, the PRM can assume a queuing model where qd(e,i) is a function of the congestion state at element e when packet i arrives.

[0051] Another illustrative network bandwidth estimation procedure includes the PathCos++ procedure. PathCos++ can use the PRM (i.e., self-congestion) to estimate the available bandwidth of a network path. For example, PathCos++ sends a periodic probe sequence. The time interval between two packets can define an instantaneous rate, and the network will react to that rate. The probe rate on the probe sequence can be decreased by increasing the time between probe packets to create a congestion state, and then gradually relieving this congestion. The PathCos++ procedure can measure the relative OWD of each probe packet in the queue, and attempts to find a pair of packets with similar OWD on either side of a congestion peak. Similar OWD can correspond to packets with similar congestion. The PathCos++ procedure can attempt to find the pair of packets with the widest gap, and then calculate the average reception rate of the probe packets between the two packets of that pair. The average reception rate can be used as an available bandwidth estimate.

[0052] Another illustrative network bandwidth estimation procedure includes the aforementioned Bump-in-the-Dark Algorithm (BDA). BDA can be integrated with PathCos++ to select a pair of packets with similar congestion on either side of a congestion peak. The selected pair of packets can be used to estimate the available bandwidth. The quality of the available bandwidth estimate can be only as good as the selection of those packets by BDA.

[0053] In BDA, the probing sequence has decreasing rates to first create a congested state of the network path (e.g., above the bottleneck rate) and then decongest the network path (e.g., below the bottleneck rate). This means that the OWD of the packets rises (congested state) and then falls (decongested state) across the probing sequence. Large spikes in the OWD can indicate the time of maximum congestion, and packets with similar OWDs should experience similar congestion (e.g., similar amount of queuing).

[0054] Another illustrative network bandwidth estimation procedure includes packet loss from errors and from congestion. For example, many networks are lossy, which means that packets can be lost or intentionally dropped. Some networks are lossless, meaning that they guarantee no packet loss. Lossless networks can only be applicable to relatively small networks because it can have large-scale adverse effects, such as head-of-line blocking. Thus, most large networks and network paths are considered lossy.

[0055] One cause of packet loss is bit errors in network devices and transmission errors over noisy links. However, these errors are not common. Most link technologies deployed in noisy channels (such as WiFi and 3G) use link-level acknowledgments and retransmissions to hide link losses. Thus, actual packet loss due to noise and errors is very rare unless conditions are poor. The reason is that TCP congestion control assumes packet loss is due to congestion, so the best way to get optimal performance is to hide link-level losses.

[0056] Thus, on lossy network paths, packet loss is almost always caused by congestion, and sustained congestion can produce packet loss. If the input rate of traffic is consistently higher than the rate at which the link can transmit, this imbalance can be resolved by dropping excess traffic and causing packet loss / drop. Thus, packet loss can be a reliable indicator of network congestion.

[0057] Another illustrative network bandwidth estimation procedure includes basic queue tail-drop loss. A link between the sender computing device 110 and the receiver computing device 120 via the network 130 can be a simple queue of fixed capacity before the link. The sender computing device 110 can transmit packets from the queue on the link at the fastest speed the link can achieve, while the sender computing device 110 can drop packets when the queue is full.

[0058] Traffic is often bursty, so queues can be used to accommodate temporary excesses in input rate while smoothing processing on the link. As long as the queue is not full, any received packet is either transmitted (if the link is idle) or added to the queue to be transmitted when the link grants. When the queue is full, any received packet can be dropped. Packet drops can occur at the tail end of the queue, so the name tail-drop loss.

[0059] With constant traffic and ideal queues, packet loss can be fine-grained and spread smoothly across the chirp sequence. In practice, however, losses can cluster together rather than being evenly distributed across the chirp sequence. A first reason can be bursty cross traffic, which changes the load on the queue between bursts, during which the queue has a chance to accept more probe packets. In contrast, the queue can accept fewer packets during bursts. A second reason can be granular queue scheduling, e.g., after a number of time slots available in the scheduling and filling queue, a number of packets are removed from the queue together, causing no time slots to open before the next scheduling. Thus, packet loss with tail-drop is typically clustered. In the traffic received at the receiver computing device 120, sequences with little or no loss alternate with sequences with very high or complete loss. This makes the congestion signal quite coarse.

[0060] Tail-drop queues can behave in two main ways. If cross traffic does not saturate the queue, then the queue is mostly empty. In this case, congestion caused by the probe / chirp sequence causes the queue to fill up, and the queuing delay increases until the queue is full, at which point the queuing delay stops increasing, and packets start being dropped. If cross traffic saturates the queue, then the queue is mostly full. In this case, congestion caused by the probe / chirp sequence can cause the queuing delay not to increase (e.g., the queue cannot become more full) and immediate packet loss. The queuing delay can fluctuate based on the burstiness of the cross traffic across the filling queue.

[0061] Another illustrative network bandwidth estimation procedure includes loss from Active Queue Management (AQM). AQM can remedy the deficiencies of tail-drop, including large queuing delay and coarse congestion signal. AQM can add extra processing in the queue, making packet loss more proportional to the degree of congestion.

[0062] AQM can implement random early detection (RED), which defines a probability of dropping / packet loss based on queue occupancy. When the queue is nearly empty, the probability of dropping a packet is close to 0, and the likelihood of packet loss / drop is very small. When the queue is nearly full, the probability of dropping a packet is higher, and thus the likelihood of packet loss / drop is greater. When a packet is received, RED can use the current probability of packet loss / drop and a random number to determine whether the packet is dropped or put into the queue. Packet loss becomes more probabilistic based on queue occupancy.

[0063] AQM can implement probability algorithms other than RED (e.g., Blue, ARED, PIE, CoDel, etc.) to maintain small / sparse occupancy of the queue by proactively dropping packets. For example, a CoDel process can attempt to keep the queue delay for all packets below a threshold. In another example, a PIE process has a burst protection feature where no packet loss / drop occurs for the first 150 ms after the queue starts to fill up.

[0064] In some examples, the short packet sequence used by the estimation technique can be too short to trigger a strong AQM response, and the loss caused by the AQM can be very low on each chirp sequence. The chirp sequence can be more likely to overflow the queue and cause tail-drop loss compared to seeing significant AQM loss. Thus, for the purpose of bandwidth estimation, the AQM queue can be considered the same as a tail-drop queue.

[0065] Another illustrative network bandwidth estimation process includes a rate limiter, policer, and token bucket. A rate limiter can be used to help network traffic conform to a certain rate. This can be for policy reasons, for example, as a result of a contractual agreement.

[0066] A rate limiter can be implemented with or without a queue. When a queue is used, the rate limiter is similar to a fixed-capacity link and can be managed by tail-drop or using AQM. When a queue is not used, the rate limiter is implemented in a simpler way and can reduce resource usage. Non-queue implementations of rate limiters are often referred to as policers or meters to distinguish them from other rate limiters. These rate limiters can be implemented using a token bucket to accommodate traffic bursts.

[0067] When implementing a token bucket, the token bucket can hold virtual tokens associated with its maximum capacity and refill tokens at a desired rate. When a packet arrives, if the bucket is not empty, the packet is passed to the queue and a token is removed. If the bucket is empty, the packet is discarded. If the link is underutilized for a period of time, the bucket can be full and excess tokens beyond the capacity will be discarded. When the link is underutilized, the token bucket capacity (or burst size) can allow for an unrestricted packet rate.

[0068] Token bucket-based rate limiters have no queue and, as a result, packets can not experience additional queuing delay due to congestion. Congestion signals can correspond to packet discards. If the size of the bucket is appropriate for the link capacity and traffic bursts (i.e., not too small nor too large), packet loss can be fairly fine-grained. In practice, configuring the burst size is tricky, so most token buckets do not produce a smooth pattern of packet loss.

[0069] Another illustrative network bandwidth estimation procedure includes PRM and packet loss. PRM assumes that no probe packets are lost on the network path and that any packets sent are received and can be used by the PRM method. Except for the NEXT-v2 procedure, existing PRM methods do not attempt to tolerate packet loss. The NEXT-v2 procedure assumes that packet loss is due to random link errors. When a packet is lost, the NEXT-v2 procedure attempts to reconstruct the packet by interpolating from neighboring packets. The NEXT-v2 procedure can tolerate limited packet loss and it can not be helpful for loss due to congestion. Other methods can assume that if there is packet loss in a chirp sequence, estimation cannot be performed. In such cases, the entire chirp sequence can be discarded and estimation is not performed. This can significantly reduce the probability of obtaining a bandwidth estimate.

[0070] When a bottleneck on a network path is based on a token bucket, PRM methods can not accurately estimate available bandwidth. For example, PRM methods measure an increase in OWD due to congestion. With a token bucket, there is no queue, and the bottleneck can not react to congestion by increasing queuing delay. In this sense, the congestion state created by PRM methods does not increase OWD. In the presence of a token bucket, PRM methods will typically fail or give erroneous estimates. As such, PRM methods can not currently provide estimates of available bandwidth when the bottleneck is a token bucket.

[0071] Examples of the present disclosure can improve these and other approaches to conventional bandwidth estimation by performing BDA bandwidth estimation filtering based on packet loss patterns. Based on identified packet loss patterns, a hypothesis or determination can be made as to whether BDA estimates can be used or assumed to be accurate. Additional actions can be performed, such as automatically re-routing packets and / or load balancing network traffic in view of acceptable BDA estimates. Implementing this process in a network can be easily implemented for an engineer or technician and can be performed with low overhead (e.g., minimizing CPU usage, fast execution so as not to slow down network path estimation, etc.).

[0072] As described above, certain types or categories of packet loss can occur, four of which are described below.

[0073] According to a first category (Category 1), transmission errors on the link overwhelm the link layer reliability mechanisms and create packet loss (i.e., packet loss from errors and congestion described above). Bit errors in the network device can cause packet loss. However, these errors are mostly random and independent of congestion, so they are not generally considered a cause of queuing delay. Since the errors are mostly random, the packet loss pattern or signature for Category 1 packet loss is fairly random. In some examples, if the probability of packet loss (relative to a given threshold) is low enough, then the BDA estimates can be considered accurate enough for use. On the other hand, if the probability is high relative to a given threshold (or maximum threshold), then the BDA estimates will be discarded.

[0074] As Figure 2 shown, the OWD (delay) fluctuations 200 across probe index / packet 5 (202a) to about probe index / packet 58 (202b) of 0 to 1500+ microseconds depict the degree / cycle of congestion reflected as OWD fluctuations, indicating an increase in congestion and subsequent decrease. These transmissions do not cause packet drops (no packet drops). In this example, the measurement points 204a / b (corresponding to packets) are selected by the BDA, which have approximately the same OWD of about 200 microseconds during the increasing and decreasing OWD.

[0075] According to a second category of packet loss (Category 2), a bottleneck can be achieved relative to a tail-drop queue or AQM queue, which is not congested by default. Thus, by default, such a queue is mostly empty and does not need to drop packets. If additional traffic is added and subsequently causes congestion, the queue will fill up, the queuing delay will increase, and packets will be dropped when the queue is full.

[0076] Figure 3 An example scenario is illustrated in which a tail-drop queue (which will be similar to an AQM mode) is used to control packet transmission and the resulting Category 2 loss. As Figure 3As shown, the OWD bump 300 occurs between about probe indices 250 and 2500 (the boundaries of the OWD bump are reflected in the points corresponding to OWD of 0 and the selected measurement points, i.e., 302a / b and 304a / b, respectively). It should be appreciated that when using a tail-drop (or AQM) queue, congestion can start to occur / increase (indicated by the increase in OWD from 0 microseconds to about 14,000 microseconds). When the queue fills up, packets will start to be dropped. In Figure 3 In the middle, this packet loss occurs between about the 500th and 2000th packet. The packet loss pattern shows this as a truncation of the OWD bump, and the selected measurement points 304a / b are "outside" the packet loss region of the OWD bump.

[0077] According to a third category (Category 3), the bottleneck can be a tail-drop queue or AQM queue, again, but by default is congested. In other words, such a queue is mostly full, and packets beyond the queue capacity will be dropped. If additional traffic is added and causes congestion, the queuing delay will not increase, but the packet loss will increase.

[0078] According to a fourth category (Category 4), the bottleneck reacts to congestion by dropping packets and not increasing the queuing delay. That is, in such a scenario, a token bucket policer can be operating in the network. If the default bottleneck is not congested, no packet loss occurs, but if congestion occurs, packets are dropped. If additional traffic is added and causes congestion, similar to Category 3, the packet loss will increase. However, it should be appreciated that such a bottleneck does not cause any increase in queuing delay. However, if there are other queues on the network path that are not the bottleneck, but still often produce an increase in delay when they are loaded. The received packets can exhibit a delay signature related to the queues that are not the bottleneck. As noted above, the PRM model assumes that no probe packets are lost on the network path, and any sent packets are received and can be used.

[0079] Figure 5 Figure illustrates an example scenario in which a token bucket policer is used to manage the transmission of packets on a network path. It can be appreciated that the OWD bump 500 is reflected at the onset of congestion, but the OWD bump is not related to the policer / primary bottleneck. It should be noted that there can be scenarios in which another queue can act as another / secondary bottleneck (as described above). Thus, the bottleneck attributed to the token bucket policer is considered the primary or main bottleneck. Most token bucket policers exhibit bursty behavior, and at the onset of congestion, it takes a period of time for the token bucket to run out of tokens and start dropping packets (causing packet loss), during which time the traffic can pass through the token bucket policer at a higher rate.

[0080] Additional queues can exist downstream of the token bucket policer, and during a burst, they can temporarily congest and create an OWD spike (approximately packet 0 to packet 100). Once the token bucket policer starts dropping packets, the downstream congestion is removed (approximately packet 100 to packet 200). After this, congestion only causes packet loss (approximately 200 to 3000 packets).

[0081] When deploying such PRM methods on real networks, especially for longer probe sequences, scenarios of congestion causing packet loss are encountered. Most packet loss is observed to be of category 2 packet loss (default non-congested AQM or tail-drop queue). To address this problem, regular PRM methods can be modified to tolerate modest packet loss. For the reduced rate method using BDA, the calculation of the average reception rate of probe packets between the two packets of a packet pair can be modified to account for lost probe packets, where the number of received packets is used instead of the number of sent packets minus the losses. These methods are generally resilient to packet loss due to congestion, as losses are unlikely to occur at the bottom of the OWD spike (less congestion). While it can be sufficient for accurate bandwidth estimation for category 2 type of packet loss, when discussing problems of other categories of packet loss, it can cause erroneous estimates. That is, losses due to tail-drop or AQM queuing methods can be safely ignored, as congestion creates enough queuing delay increase to create an OWD spike that can be used for bandwidth estimation.

[0082] However, for other categories, bandwidth estimation can be inaccurate. For example, for category 1 packet loss, if those packet losses occur after the bottleneck, these packets are subtracted from the reception rate, and can cause underestimation. For category 3 type of packet loss, queuing delay is not increased, no OWD spike is created, and therefore, bandwidth estimation will fail, i.e., such packet losses are not correctly considered to affect the available bandwidth on the network path. For category 4 type of packet loss, the main bottleneck does not increase queuing delay, and no OWD spike is created. Therefore, either no OWD spike is created, and therefore bandwidth estimation will fail, or the OWD spike is created by a queue that is not the bottleneck, and the resulting bandwidth estimation will overestimate the available bandwidth.

[0083] Examples of the present disclosure can improve bandwidth estimation by taking into account the rate of packet loss, and only perform bandwidth estimation when the packet loss across a probe sequence is moderate (e.g., less than 10% or 20%), while discarding the bandwidth estimation of the BDA in other cases. However, such filtering is not sufficient. For example, in many scenarios, a Category 1 or Category 4 type of situation can cause moderate packet loss, leading to an erroneous bandwidth estimation. There are also situations where heavy packet loss occurs due to control using tail-drop or AQM queues, and where estimation can be performed using BDA.

[0084] Examples of the present disclosure are directed to LBF, where packet loss associated with a particular type(s) of packet loss category can be processed, but can do so without expensive computation and without adding estimation error to the available bandwidth estimation. Examples of the present disclosure can analyze the packet loss pattern to determine whether packet loss (dropped packets) occurred within the OWD spike. If so, such packet loss can be assumed to be comparable to a Category 2-type of packet loss, and as will be described below, the BDA estimation can be assumed to be sufficiently accurate / precise for network traffic engineering or other applications / contexts that need to know the available bandwidth on a network path. If not, i.e., the packet loss pattern is associated with another category type, then the BDA can be assumed to produce an inaccurate or erroneous available bandwidth estimation, and that estimation can be discarded, ignored, or otherwise not used.

[0085] In other words, it can be determined whether the BDA estimation is acceptable in the presence of packet loss on a probe sequence, allowing packet loss to be integrated (factored in) into PRM bandwidth estimation techniques. In particular, only minor adjustments to conventional BDA are needed when used in popular estimation methods, to better determine when BDA can be used for a probe sequence with packet loss. Thus, examples of the present disclosure can be readily implemented, and CPU usage and other overhead can be minimized / minimalized. This helps to avoid slowdowns due to network path BDA estimation.

[0086] Referring back to Figure 2 When a probe sequence used for BDA-based estimation is properly configured, a congestion spike occurs, followed by a drop / decongestion. In the absence of any packet loss, such a scenario is characterized by an OWD spike, again, the delay rises and returns to the uncongested delay. After the OWD spike, the delay pattern (more or less) tracks the minimum OWD threshold. Referring back to Figure 3In the case of packet loss due to the use of tail-drop or AQM queues, it can be appreciated that the OWD hump is identical / similar to the case without packet loss except that the upper portion of the OWD hump is truncated. Thus, when the tail-drop or AQM queue is full, the delay is capped rather than rising to a particular level / value. As noted above, packet loss occurs when the queue is full, which is the only time that packets are dropped, which is consistent with high delay, a typical characteristic of such Class 2 type scenarios is that packet loss occurs in the middle portion of the OWD hump, i.e., the peak of congestion.

[0087] Other categories of packet loss described above do not exhibit this unique signature. Transmission loss (Class 1) is random and uniformly distributed across the entire chirp sequence independent of congestion. In the case of a queue that has saturated (Class 3), the queue is always congested and thus packet loss occurs throughout the probe sequence. When token bucket policing is used (Class 4), packet loss is expected to occur until the probe rate is below the rate of the token bucket, where the token bucket does not result in an increase in delay. If there is an OWD hump associated with another queue (a downstream queue that is temporarily congested), as described above, then this queue is not the cause of the primary bottleneck and has a higher rate than the ABW. Thus, the OWD hump terminates before the ABW, and packet loss can occur beyond the termination point of the OWD hump. From Figure 5 It can be appreciated from

[0088] As noted previously, the BDA detects the OWD hump in the received probe sequence, selects two packets in the OWD hump, and then performs a measurement between the two packets, which can be used to infer an estimate of the available bandwidth. The various BDAs can differ in the manner in which they perform these tasks (e.g., PathCos++, SLDRT, Voyager-2, Voyager3 can use different methods to select the measurement points / packets). However, most BDAs attempt to find two packets on either side of the OWD hump, near the bottom of the OWD hump (i.e., near the beginning of the rise in congestion and near the end of the decongestion before returning to normal). Typically, the BDA does not attempt to perform a measurement across the entire OWD hump. Rather, most BDAs simply identify two measurement points, i.e., packets in the probe sequence that have the same / similar level of congestion, which are used to determine the available bandwidth estimate.

[0089] As discussed previously, packet loss is expected to occur at the top of the OWD bump when tail-drop or AQM queue is not yet saturated. However, due to variations in delay and loss batching, the top of the OWD bump can not be well-defined in various scenarios, and packet loss can occur during low queue utilization, i.e., when the OWD is still relatively small. Some queues process packets in batches, causing packet loss to bunch together, and the OWD at the top of the OWD bump has a jagged or non-flat signature. Referring now to Figure 6 , an example scenario of packet drops caused by tail-drop and AQM in a congested queue is illustrated. It should be understood that the goal of AQM is to start dropping packets before the queue is full. In this way, TCP senders can detect congestion earlier. When used with TCP, AQM allows TCP to react faster and avoid filling up the entire queue. In Figure 6 , packet loss due to tail-drop occurs at the top of the OWD bump 600 between packets 500 and 3000. Packet loss due to AQM occurs when the queue is filled up (delay increases) between packets 100 and 500, and when the queue shrinks (delay decreases) between packets 3000 and 3800. It can be appreciated that packet loss starts before the OWD reaches the top of the OWD bump 600, and continues after the OWD starts to decrease. Thus, the packet loss region occurs approximately between packets 200 to 3800, and the top of the OWD bump 600 is not strictly flat because AQM reduces the queue size. It should be understood that the "top" of the OWD bump can refer to the portion(s) of the OWD bump where the packet OWD is close to the maximum OWD (e.g., the difference from the maximum OWD is less than about 10% of the difference between the maximum and minimum OWD). This is in contrast to a strict tail-drop queue, where packet loss only occurs at the top of the OWD bump ( Figure 3 ), although AQM is not necessarily the only reason for loss to occur outside the top of the OWD bump. For example, other network elements on a particular network path can have variable delays, adding noise to any OWD measurement.

[0090] As mentioned above, BDA assumes that when a queue has processed a packet selected as a measurement point, the delay experienced by that packet properly reflects the congestion level. If packet loss does occur, the delay does not indicate full congestion. Thus, for accurate measurements, there should not be packet loss around those selected packets / measurement points (at least, no significant packet loss). If there is packet loss in the portion of the OWD bump between measurement points, it can be assumed that they are primarily caused by unsaturated tail-drop, and can be safely ignored as long as those packet losses are not close to the measurement points. Another aspect, around or outside of those packets (e.g., in the portion of the OWD bump between measurement points), the delay can be assumed to be dominated by the queue size, and thus can be used to measure the congestion level. Figure 3Any significant packet loss occurring (to the left of the first measurement point / packet 302a and to the right of the second measurement point / packet 302b) indicates that the packet loss is due to another cause and the bandwidth estimate is inaccurate and should not be used.

[0091] The region of negligible packet loss can be defined based on the selected / packets for measurement Figure 4 ). As the packets for measurement are the consideration for obtaining an accurate BDA estimate, this region can include a number of packets before / after the actual measurement point / packet where packet loss occurrence is acceptable, in other words, such packet loss can be ignored. Typically, one to three packets before / after the actual measurement point or packet are sufficient to capture the occurrence of acceptable (or unacceptable) packet loss, although the actual number of packets before / after the measurement point / packet can vary, for example, based on a given percentage of the increase / decrease portion or zone / region of the OWD spike. That is, to ensure (or attempt to ensure as much as possible) that the selected pair of packets (measurement points) are not affected by packet loss, where ideally, there should be no, or in some cases, at least very few, packet loss around / near the packet pair. According to some examples of the disclosed technology, the region of negligible packet loss excludes several packets around the packet pair. This number of packets can be statically configured, it can be a fraction of the total number of packets in the chirp sequence, or it can be based on other measurements performed by the receiver. For example, if this number of packets is two, the region of negligible packet loss extends from two packets after the first measurement point / packet to two packets before the second measurement point / packet. The effect is that if there is significant packet loss in the initial portion of the probe sequence (up to 2 packets after the first measurement packet) and in the last portion of the probe sequence (starting 2 packets before the second measurement packet), the resulting bandwidth estimate can be discarded.

[0092] Figure 4 Such examples are illustrated in FIG. 4, where the dashed lines 402a / b corresponding to (about) packets 400 and 2200 depict the extent of the region of negligible packet loss based on the selected packets for measurement. In Figure 4In some examples, the negligible packet loss region can be defined by a threshold value, which can be defined as a percentage of the total number of packets in the probe sequence. For example, if the probe sequence includes 100 packets, then a threshold value of 10% can be defined. In this case, the negligible packet loss region can be defined as the region where the packet loss is less than 10% of the total number of packets. In other words, the region between the two dashed lines 402a / b can be defined as the region where the packet loss is less than 10% of the total number of packets in the probe sequence. In this case, the packet loss 300 caused by tail drop is completely between the two dashed lines 402a / b, and thus in the region of negligible packet loss, and does not cause the bandwidth estimate to be discarded.

[0093] In other examples, alternative definitions of the negligible packet loss region can include, but are not limited to, for example, defining the top of the OWD bump based on the delay, or even using the full range / width of the OWD bump to determine whether the packet loss falls within an acceptable range. In some examples, the negligible region of packet loss can be defined by two packets whose associated OWD delays occur approximately halfway between the bottom (lowest delay value) and the top / peak (highest delay value) of the OWD bump. Figure 4 Such examples are illustrated in Figure 4, where the lines 400a / b can define the range of the negligible packet loss region, and can be selected based on the packets that are halfway between the minimum and maximum OWD. As can be appreciated, the packet loss 300 caused by tail drop is completely between the lines 400a / b, and thus in the region of negligible packet loss.

[0094] It will be appreciated that, in general, due to the noise / burst distribution of losses and transmission errors (category 1), it cannot be assumed that the parts of the probe sequence outside the negligible packet loss region (400a / b or 402a / b) will be completely free of loss. Therefore, determining whether the BDA estimate can be used should be sufficiently robust to take into account real-world conditions (i.e. the nature of the noise and bursts). In some examples, such robustness is achieved by using a packet loss percentage threshold, which can be used to account for such characteristics of the probe chain / network.

[0095] In particular, if the percentage of packet loss outside the negligible packet loss region is a relatively low value, e.g., below a user-defined threshold by some determined amount, e.g., the default 5%, it can be assumed that the packet loss occurring before or after the OWD bump is due to noise / transmission loss (category 1) and not “real” packet loss due to another packet loss category. Thus, when the packet loss is below this threshold, the BDA estimates can be used. Such a threshold is typically much lower than the packet loss threshold applied to the entire probe chain. This more stringent threshold is possible because the packet loss is expected to be within the bounds of the OWD bump, and not at or outside the measurement point(s) / packets where packet loss is expected to be lower. Thus, in some examples, more BDA estimates from the probe sequence can be discarded when the packet loss is due to other conditions / causes. This more stringent filtering in turn can result in more accurate BDA estimates.

[0096] Figure 7 and Figure 8 A system and corresponding method of operation for bandwidth estimation filtering are illustrated according to one example of the disclosed technology, and will be described in conjunction with the drawings. As shown, a recipient computing device 720, such as a network device or element, can receive a set of probe packets (a set of packets) from a network 730 over a network path (similar to the scenario shown in Figure 7 Figure 1 As shown, the recipient computing device 720 can receive a set of probe packets (a set of packets) from a network 730 over a network path (similar to the scenario shown in

[0097] The recipient computing device 720 can include one or more computing components implemented by a hardware processor 722, e.g., one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for retrieval and execution of instructions stored in a machine-readable storage medium 724. The hardware processor 722 can fetch, decode, and execute instructions, such as instructions to implement the methods / operations shown in Figure 8 As an alternative or in addition to retrieving and executing instructions, the hardware processor 722 can include one or more electronic circuits comprising electronic components for performing the functionality of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.

[0098] ​A machine-readable storage medium, such as the machine-readable storage medium 724, can be any electronic, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, the machine-readable storage medium 724 can be, for example, Random Access Memory (RAM), non-volatile RAM (NVRAM), Electronically Erasable Programmable Read-Only Memory (EEPROM), a storage device, an optical disc, and the like. In some examples, the machine-readable storage medium 724 can be a non-transitory storage medium, where the term "non-transitory" does not encompass transitory propagating signals. As discussed in detail below, the machine-readable storage medium 724 can be encoded with executable instructions.

[0099] In some examples, the recipient computing device 720 can include a packet receiving component 720a. As shown, at operation 800, a set of probe packets is received over a network path. The set of probe packets can be a sequence of data packets including a probe sequence as described above, with a particular time interval between each probe packet departure. The time interval between packets defines a momentary rate at which the network can react by keeping the time interval substantially constant or spacing out the packets. The probe packets can be sent in order to inject delay / congestion on the network path, so that bandwidth estimation can be performed. Figure 8

[0100] In some examples, the recipient computing device 720 can include a bandwidth estimation component 720b. Delays are measured over the network path, and can be input into the bandwidth estimation component 720b. As discussed previously, the delays associated with the received packets can be measured using the OWD process to determine the network path bandwidth. Additionally, any packet loss is also measured, and can be input into the LBF component 720c (described below).

[0101] For example, and as shown, at operation 802, the available bandwidth of the network path is estimated using BDA. Again, BDA involves selecting a pair of packets with similar congestion on both sides of a congestion peak or OWD bump. The pair of packets selected to be used as a measurement point can be used to estimate the available bandwidth in the presence of a bottleneck that reacts to congestion state by dropping packets (i.e., creating packet loss). In BDA, the probe sequence has a decreasing rate to first create a congestion state of the network path (e.g., a rate above the bottleneck) and then decongest the network path (e.g., a rate below the bottleneck). This means that across the probe sequence, the OWD of the packets rises (congestion state) and then falls (decongestion state). When a bottleneck occurs due to the use of a tail-drop queue or AQM (which is not congested by default), additional traffic / increased traffic causes congestion. The queue will fill up, the queuing delay will increase, and packets will be dropped when the queue is full. Figure 8 ​​

[0102] Some margin or allowance for error is provided due to noise, burstiness of packet transmission, and transmission errors. In some examples, this error tolerance can be provided by the LBF component 720c, which can determine whether there is a negligible region of packet loss 720 around the measurement point / packet pair. As shown, at operation 804, a negligible region of packet loss is determined with respect to the OWD spike reflecting transmission delay / congestion of the set of probe packets on the network path. In some examples, a particular threshold can be set with respect to the region of negligible packet loss. When packets are dropped at a reasonable rate (e.g., falls within a given threshold or is otherwise acceptable), it can be assumed that the dropped packets occurred due to tail drop or AQM queue becoming full, in which case the bandwidth estimate derived using BDA ignoring packet loss can be deemed accurate. That is, at operation 806, upon determining that the amount of dropped packets is sufficiently low outside the region of negligible packet loss, the estimated available bandwidth can be accepted for some subsequent action(s), e.g., such as an action associated with network traffic engineering. In other words, the available bandwidth estimate can be reported to the network traffic engineering performance component 720d. When packets are dropped at an unreasonable rate (e.g., falls outside a given threshold or is otherwise unacceptable), it can be assumed that the dropped packets occurred due to some reason(s) other than tail drop or AQM queue becoming full, in which case the bandwidth estimate derived using BDA ignoring packet loss can not be deemed accurate. That is, the packet loss cannot be reasonably ignored. Accordingly, at operation 808, upon determining that the negligible region of packet loss reflects an unacceptable amount of dropped packets, the estimated available bandwidth can be discarded or otherwise suppressed. In some examples, another BDA-based estimation(s) can be performed, where another set of probe packets are received and analyzed as described herein. At operation 810, the estimation of available bandwidth can be reported out and / or used to perform network traffic engineering actions. Figure 8

[0103] ​In some examples, the ignorable region is related to a certain subset of packets of a probe sequence that are proximate to falling within a certain acceptance threshold level of a measurement point / packet. For example, an ignorable region of two packets can be specified or selected, where the two packets are based on the packets used for measurement by the BDA. In other examples, the two packets can be the packets bounded by the measurement point / packet. In other examples, the two packets can be specified based on the maximum latency of the OWD bulge identified by the BDA. In other examples, the two packets can be specified based on the entirety of the OWD bulge. The threshold applied can vary, but in some examples, the threshold can be calculated based on the maximum and minimum amount of latency that occurs within the OWD bulge. In some examples, the threshold can be a strict / absolute filter, where packet loss proximate to the measurement point / packet is completely unacceptable, meaning that all packet loss must occur within the OWD bulge in order to be acceptable for such bandwidth estimation.

[0104] It should be appreciated that the examples described herein are premised on ignoring packet loss in the OWD bulge. One benefit of ignoring packet loss in the OWD bulge is that doing so enables bandwidth estimation in more conditions or scenarios. Further, ignoring packet loss in the OWD bulge allows for determining (and using) bandwidth estimations that were not previously available and accurate enough. Further, since a stricter threshold can be implemented to eliminate or ignore probe sequences with unwanted packet loss, the accuracy of the bandwidth estimations obtained by looking for specific class 2 type bottleneck or path loss patterns as described herein can be improved. Ultimately, better available bandwidth estimations improve network traffic engineering efficiency, especially in the context of SD-WAN. Another benefit is that ignoring packet loss in the disclosed manner is a non-complex modification to known (sometimes best) methods for measuring available bandwidth. Examples of the disclosed technology can be improved or modified for use in most network implementations without having to re-architect code or change packet formats.

[0105] Figure 9 A block diagram of an example computer system 900 in which various examples described herein can be implemented is depicted. For example, Figure 1 The sender computing device 110 or the recipient computing device 120 shown can be implemented as Figure 9 The example computer system 900 in FIG. 1. The computer system 900 includes a bus 902 or other communication mechanism for communicating information, and a one or more hardware processors 904 coupled with the bus 902 for processing information. The hardware processor(s) 904 can be, for example, one or more general purpose microprocessors.

[0106] The computer system 900 also includes a main memory 906, such as a random access memory (RAM), cache and / or other dynamic storage devices, coupled to bus 902 for storing information and instructions to be executed by processor 904. Main memory 906 also can be used for storing temporary variables or other intermediate information during execution of instructions by processor 904. Such instructions can make computer system 900 into a special purpose machine specific to perform operations specified in the instructions when stored in a storage medium accessible to the processor 904.

[0107] The computer system 900 further includes a read only memory (ROM) 908 or other static storage device coupled to bus 902 for storing static information and instructions for processor 904. A storage device 910, such as a magnetic disk, optical disk, or USB thumb drive (flash drive), etc., is provided and coupled to bus 902 for storing information and instructions.

[0108] Computer system 900 can be coupled via bus 902 to a display 912, such as a liquid crystal display (LCD) (or touch screen), for displaying information to a computer user. An input device 914, including alphanumeric and other keys, is coupled to bus 902 for communicating information and command selections to processor 904. Another type of user input device is cursor control 916, such as a mouse, a trackball or cursor direction keys for communicating direction information and command selections to processor 904 and for

[0109] The computing system 900 can include a user interface module implementing a GUI that can be stored in a mass memory as executable software code to be executed by a computing device. As examples, this module and other modules can include components, such as software components, object-oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.

[0110] In general, the words "component," "engine," "system," "database," "data store," and the like, can refer to logic embodied in hardware or firmware, or to a collection of software instructions, possibly having entry and exit points, written in a programming language, such as, for example, Java, C or C++. A software component can be compiled and linked into an executable program, installed in a dynamic link library, or can be written in an interpreted programming language such as, for example, BASIC, Perl, or Python. It will be appreciated that software components can be callable from other components or from themselves, and / or can be invoked in response to detected events or interrupts. Software components configured for execution on computing devices can be provided on a computer readable medium, such as a compact diskette, a digital video disk, flash drive, a magnetic disk, or any other tangible medium, or as a digital download (and can be originally stored in a compressed or installable format that is decompressed or installed at the time of execution). Such software code can be stored, for example, in the storage device components 910, and loaded and executed by the hardware processor components 904. As another example, hardware components can include connected logic elements, such as gates and flip-flops, and / or can include programmable logic elements, such as PLCs or processors. The computing system 900 can implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 900 to be a special-purpose machine. According to one example, the techniques herein are performed by computer system 900 in response to the

[0111] The computer system 900 can implement the techniques described herein using customized hard-wired logic, one or more ASICs or FPGAs, firmware and / or program logic which in combination with the computer system causes or programs computer system 900 to be a special-purpose machine. According to one example, the techniques herein are performed by computer system 900 in response to the

[0112] As used herein, the term “non-transitory medium” and similar terms, refer to any medium that stores the data and / or instructions that cause a machine to operate in a specific manner. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 910. Volatile media includes dynamic memory, such as the main memory 906. Common forms of non-transitory media include, for example, a floppy disk, a flexible disk, a hard disk, a solid state drive, a magnetic tape, or any other magnetic data storage medium, a CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, a RAM, a PROM, and EPROM, a FLASH-EPROM, NVRAM, any other memory chip or cartridge, and networked versions of the same.

[0113] Non-transitory media differ from transmission media, which are involved with transferring information from one place to another. Transmission media include coaxial cables, copper wire, and fiber optic cables, including the wires that comprise bus 902. Transmission media can also take the form of acoustic or light waves, such as those generated during radio and infrared data communications.

[0114] Computer system 900 also includes a network interface 918 coupled to bus 902. Network interface 918 provides a two-way data communication coupling to one or more network links that are connected to one or more local networks. For example, network interface 918 can be an integrated services digital network (ISDN) card, cable modem, satellite modem, or a modem to provide a data communication connection to a corresponding type of telephone line. As another example, network interface 918 can be a local area network (LAN) card to provide a data communication connection to a compatible LAN (or WAN component to communicate with a WAN). Wireless links can also be implemented. In any such implementation, network interface 918 sends and receives electrical, electromagnetic or optical signals that carry digital data streams representing various types of information.

[0115] Network links typically provide data communication through one or more networks to other data devices. For example, a network link can provide a connection through a local network to a host computer or data equipment operated by an Internet Service Provider (ISP) to other data devices operated by users. The ISP in turn provides data communication services through the world wide packet data communication network now commonly referred to as the “Internet.” Local networks and the Internet both use electrical, electromagnetic or optical signals that carry digital data streams. The signals through the various networks and the signals on network link and through network interface 918, which carry the digital data to and from computer system 900, are example forms of transmission media for digital data.

[0116] Computer system 900 can send messages and receive data, including program code, through the network(s), network link(s), and network interface(s) 918. In the Internet example, a server might transmit a requested code for an application program through the Internet, ISP, local network and network interface 918.

[0117] The received code can be executed by processor 904 as it is received, and / or stored in storage device 910, or other non-volatile storage for later execution.

[0118] Each of the processes, methods, and algorithms described in the preceding sections can be embodied in, and fully or partially automated by, code components of one or more computer programs. The one or more computer programs can be executed by one or more computer systems or computer processors comprising computer hardware, and can be implemented in an operating system or utility application

[0119] As used herein, a circuit can be implemented using any form of hardware, software, or combinations thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logical components, software routines or other mechanisms might be implemented to make up a circuit. In implementations, various circuits described herein can be implemented within the same circuit, or various circuits described can be implemented on different circuits within the same computing device or across multiple computing devices. Even if various features or elements of functionality are described or claimed as being part of a single circuit, these features and functionality can be implemented at different times or in different systems, and the various features and functionality can be implemented alone or in combination with yet other features and functionality, and such description or depiction of various order of operation or elements within a process is not necessarily meant to limit the scope of claims to that precise order or configuration. Even if particular features or elements of functionality are limited to a single circuit or implementation, those need not be the case throughout the disclosure, and the different functionality described can be implemented across multiple circuits or across a single circuit. Similarly, the various features and functionality described or indicated as being part of a single implementation can be implemented alone or in combination in various implementations. The disclosure is not limited to the described or depicted order or configuration of these operations. The various operations can be carried out in any order, depending on the implementation desired. Further, those of ordinary skill in the art will recognize that, in various implementations, the means or components for taking two or more described or claimed features can be implemented using electronic hardware, computer software, or any combination thereof. Those of ordinary skill in the art will also recognize the interchangeability of hardware and software

[0120] As used herein, the term "or" can be construed in either an inclusive or exclusive sense. Furthermore, the description of resources, operations, or structures as singular or multiple is not to be interpreted as excluding the plural or singular form, respectively. Conditional language, such as "can," "could," "might," or "may," for example, unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain examples include certain features, elements, and / or steps, while other examples do not.

[0121] Unless otherwise expressly stated, terms and phrases as used herein, and variations thereof, should not be construed as limiting. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms of similar meaning should not be construed as limiting the item described to a given time period or to an item available as of a given time, but instead should be read to encompass conventional, traditional, normal, or standard technologies that can be available or known now or at any time in the future. In some cases, the existence of broadening words and phrases such as "one or more," "at least," "but not limited to," or other like phrases can not be read to mean that the narrower case is alternatively excluded, unless expressly stated otherwise or in used in the context as understood by those of ordinary skill.

Claims

1. A method for bandwidth estimation filtering, comprising: receiving a set of probe packets over a network path; estimating, by a system comprising a hardware processor, available bandwidth of the network path using a bump in the road algorithm (BDA); identifying, by the system, a delay bump experienced by a first subset of probe packets of the set of probe packets over the network path, the first subset of probe packets comprising a series of probe packets between a pair of packets of the BDA, wherein a transmission delay of the delay bump experienced by the first subset of probe packets is greater than a transmission delay experienced by a remainder of the set of probe packets; identifying, by the system, a negligible region based on the delay bump, the negligible region comprising the first subset of probe packets, wherein packet loss of probe packets in the negligible region is ignored in estimating the available bandwidth; determining, by the system, an amount of lost probe packets associated with probe packets outside of the negligible region; determining, by the system, whether the amount of lost probe packets violates a threshold value; and based on determining that the amount of lost probe packets violates the threshold value, suppressing, by the system, use of the estimated available bandwidth in directing data traffic along the network path.

2. The method of claim 1, wherein estimating the available bandwidth using the BDA comprises: selecting a first probe packet of the set of probe packets on a first side of the delay bump and selecting a second probe packet of the set of probe packets on a second, different side of the delay bump, the first and second probe packets forming the pair of packets of the BDA, and calculating the estimated available bandwidth based on the first and second probe packets.

3. The method of claim 2, wherein the negligible region is identified based on the first and second probe packets. determining the negligible region to begin at a first specified number of probe packets after the first probe packet and to end at a second specified number of probe packets before the second probe packet.

4. The method of claim 3, comprising:

5. The method of claim 4, wherein the first and second specified numbers of probe packets are based on a configured fraction of a number of probe packets in the first subset of probe packets that experience the delay bump.

6. The method of claim 3, wherein the negligible region is identified based on a ratio between a minimum one-way delay and a maximum one-way delay within the delay bump. determining the negligible region to begin at the first probe packet at which a one-way delay is greater than the ratio and to end at the second probe packet at which a one-way delay is greater than the ratio.

7. The method of claim 6, comprising:

8. The method of claim 1, wherein the amount of lost probe packets violates the threshold value if the amount of lost probe packets exceeds the threshold value.

9. The method of claim 1, wherein packet loss experienced by the first subset of probe packets is ignored for estimating the available bandwidth. ​ 10. The method of claim 1, further comprising: using the estimated available bandwidth in directing data traffic along a network path based on a determination that the amount of lost probe packets does not violate the threshold.

11. A system for bandwidth estimation filtering, comprising: a hardware processor; and a non-transitory storage medium storing instructions executable on the hardware processor to: receive a set of probe packets over a network path; measure transmission delays experienced by probe packets of the set of probe packets; estimate an available bandwidth of the network path based on the measured transmission delays using a bump in the road algorithm (BDA); identify a delay bump experienced by a first subset of probe packets of the set of probe packets over the network path, the first subset of probe packets comprising a series of probe packets between a pair of packets of the BDA, wherein the delay bump experienced by the first subset of probe packets has a transmission delay that is greater than transmission delays experienced by a remainder of the set of probe packets; identify an ignorable region based on the delay bump, the ignorable region comprising the first subset of probe packets, wherein packet loss of probe packets in the ignorable region is ignored in estimating the available bandwidth; determine an amount of lost probe packets associated with probe packets outside of the ignorable region; determine whether the amount of lost probe packets violates a threshold; and inhibit use of the estimated available bandwidth in directing data traffic along a network path based on a determination that the amount of lost probe packets violates the threshold.

12. The system of claim 11, wherein the measured transmission delays comprise one-way delays.

13. The system of claim 11, wherein estimating the available bandwidth comprises: selecting a first probe packet of the set of probe packets on a first side of the delay bump and selecting a second probe packet of the set of probe packets on a second, different side of the delay bump, the first and second probe packets forming the pair of packets of the BDA, and calculating the estimated available bandwidth based on the first and second probe packets.

14. The system of claim 13, wherein the instructions are executable on the hardware processor to determine the ignorable region to begin at a first specified number of probe packets after a first probe packet and to end at a second specified number of probe packets before a second probe packet.

15. The system of claim 14, wherein the first and second specified numbers of probe packets are based on a configured fraction of a number of probe packets in the first subset of probe packets that experience the delay bump.

16. The system of claim 13, wherein the ignorable region is identified based on a ratio between a minimum one-way delay and a maximum one-way delay within the delay bump. ​ 17. The system of claim 16, wherein the instructions are executable on the hardware processor to determine the negligible region to start at the first probe packet for which one-way delay is greater than the ratio and to end at the second probe packet for which one-way delay is greater than the ratio.

18. The system of claim 11, wherein the instructions are executable on the hardware processor to: based on determining that the amount of lost probe packets does not violate the threshold, use the estimated available bandwidth in directing data traffic along a network path.

19. A non-transitory machine-readable storage medium comprising instructions that, when executed, cause a system to: receive a set of probe packets over a network path; measure transmission delays experienced by probe packets in the set of probe packets; based on the measured transmission delays, estimate an available bandwidth of the network path using a bump in the road algorithm (BDA); identify a delay bump experienced by a first subset of probe packets in the set of probe packets over the network path, the first subset of probe packets comprising a series of probe packets between pairs of packets in the BDA, wherein the delay bump experienced by the first subset of probe packets has transmission delays that are greater than transmission delays experienced by a remainder of the set of probe packets; based on the delay bump, identify a negligible region, the negligible region comprising the first subset of probe packets, wherein packet loss of probe packets in the negligible region is ignored in estimating the available bandwidth; determine an amount of lost probe packets associated with probe packets outside of the negligible region; determine whether the amount of lost probe packets violates a threshold; and based on determining that the amount of lost probe packets violates the threshold, suppress use of the estimated available bandwidth in directing data traffic along a network path.

20. The non-transitory machine-readable storage medium of claim 19, wherein the delay bump is identified by the BDA.

Citation Information

Patent Citations

  • Method and a device for determining bandwidth transmission capability

    CN109905257A

  • Available network bandwidth estimation by using one-way-delay noise filter with bump detection

    CN112889243A