Communication problem causality identification
Patent Information
- Application Number
- CN202510928373.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2025-02-18
- Filing Date
- 2025-07-07
- Publication Date
- 2026-08-18
Smart Images

Figure CN122601440A_ABST
Abstract
Description
Background Technology
[0001] As computer networks become increasingly complex, network communication can involve more and more different components. In fact, many intermediate components facilitate parts of the network communication between a client device and a destination endpoint. For example, several different data forwarding devices (such as routers, switches, etc.) can forward data from a client device to the destination endpoint and vice versa. Furthermore, some of these intermediate components can reside in a local network (e.g., a local area network (LAN)), while others can reside in a remote or non-local network (e.g., a wide area network (WAN)). Attached Figure Description
[0002] The features, aspects, and advantages of this disclosure will be better understood when the following detailed description is read with reference to the accompanying drawings, in which similar characters denote similar parts, wherein:
[0003] Figure 1 This is a block diagram illustrating a network-based communication system with a network problem causality detection system according to various aspects of this disclosure;
[0004] Figure 2 This is a block diagram illustrating a system for providing a network path between a client device and an endpoint according to various aspects of this disclosure, wherein cause-and-effect detection of network problems is implemented;
[0005] Figure 3 It is a flowchart depicting a process for identifying the cause of communication problems in a communication session between a client device and an endpoint, according to various aspects of this disclosure;
[0006] Figure 4 It is a flowchart depicting a process for identifying an endpoint as the cause of a communication problem using performance problem indicator (IPI) data, according to various aspects of this disclosure;
[0007] Figure 5 It is a flowchart depicting a process for combining IPI data with queue monitoring data to determine the causation of network congestion problems according to various aspects of this disclosure; and
[0008] Figure 6 This is another flowchart depicting a process, according to various aspects of this disclosure, for combining IPI data with queue monitoring data to determine the cause of network latency in communication problems. Detailed Implementation
[0009] As network infrastructure becomes increasingly complex and extensive, identifying and locating the causes of network communication problems has become difficult and resource-intensive. Analyzing packet drops across network components can provide an indication that a specific network component is experiencing overload. However, such analysis of packet drops may not provide a fine-grained identification of the cause of communication problems during a specific communication session between a client device and an endpoint, potentially leading to incorrect attribution of network communication problems to the wrong network component and / or network location.
[0010] For example, packet dropping may not explain other causes of communication problems. In fact, packet dropping may not provide enough information to determine whether the communication problem is occurring at an endpoint (e.g., the destination endpoint of network communication from a client device), at the network service provider (e.g., the WAN operator), or at another source. Therefore, analyzing packet dropping in isolation may provide an incomplete analysis of the communication problem and, in some cases, may provide a false indication of the causality of the network communication problem.
[0011] Furthermore, intermediate routing devices are often shared. Numerous packets associated with different communication sessions between various client devices and endpoints can be delivered through a common intermediate routing device at any given time. Therefore, identifying a specific intermediate routing device experiencing packet dropping does not resolve network issues related to a particular communication session. Consequently, when a specific client device encounters a communication problem, packet dropping analysis itself may not be related to that particular client's network session and therefore may not be the cause of the communication problem for that specific client device. Similarly, it can be difficult to determine whether other communication sessions are burdening certain intermediate routing devices (e.g., causing switches to drop packets) and thus causing communication problems across the shared network. Therefore, it can be difficult to diagnose the actual network components causing the communication problem for a specific client device and to determine appropriate mitigation actions, such as performing load balancing or initiating other mitigation techniques to resolve the communication problem encountered by the client device.
[0012] Furthermore, as networks evolve, more and more packets are transmitted through an increasing number of switches, potentially making the resources available for analyzing packet drops enormous. A single client can transmit large numbers (e.g., thousands and / or millions) of packets across a wide variety of intermediate routing devices to facilitate a single communication session. Large-scale network monitoring (especially considering the number of client devices that can communicate across the network simultaneously) can be a significant drain on both hardware and time resources.
[0013] In light of the foregoing, this disclosure describes techniques for effectively identifying the location (e.g., network components and / or specific networks (e.g., LAN or WAN)) of network communication problems between client devices and endpoints. More specifically, this disclosure describes a workflow for analyzing performance problem indicator (IPI) data (such as Transmission Control Protocol (TCP) performance metrics) in conjunction with queue monitoring data to pinpoint components of the network path associated with the communication problem. That is, by aggregating and analyzing IPI data and queue monitoring data received from intermediate routing devices that form part of the network path between the client device and the endpoint, specific components of the network path and / or specific networks on the network path of the network communication can be identified as the cause of the network communication problem. For example, the workflow presented herein evaluates TCPIPI data obtained from access network components (such as access switches) of the local network used by the client device to access network communication, as well as queue monitoring data from multiple intermediate routing devices in the network path of the network communication, to provide an effective technique for identifying the location of a communication problem, thereby enabling more effective troubleshooting and / or mitigation responses when such a communication problem occurs.
[0014] In TCP communication, when a client device initiates a communication session with an endpoint, the client device can first communicatively connect to that endpoint across the network (referred to herein as a "TCP handshake"). TCPIPI data can be provided as part of this TCP handshake and can indicate the performance of the communication between the client device and the endpoint. Network Problem Causality Detection (NICD) systems can receive and analyze the received TCPIPI data (e.g., patterns of TCP performance metrics) to identify possible categories of communication problems. Possible categories can indicate the location of the communication problem and / or provide initial steps to locate network components that may be the cause of the communication problem. For example, as will be described in detail below, a NICD system can use IPI data to determine if the communication problem is associated with network congestion, network latency, and / or to determine if the communication problem occurs at the endpoint of the network communication. In some cases, queue monitoring data can be used to further pinpoint the location of the communication problem. That is, a NICD system can also receive queue monitoring data from multiple intermediate routing devices (e.g., access switches, aggregation switches, core switches, and WAN edge devices) in the network path between the client and the endpoint. After receiving and analyzing IPI data, the NICD system can analyze queue monitoring data at intermediate routing devices in the network path to determine the location of potential communication problems (e.g., a specific network (e.g., a LAN or WAN) and / or network components). For example, network components that may cause communication problems could be components of the network path, such as one or more intermediate routing devices, endpoints, or network service providers (e.g., a service provider offering a remote network (e.g., a WAN)).
[0015] After identifying the location (e.g., a specific network and / or network component) that may be the cause of a communication problem, the NICD system can generate an alert to network entities associated with the identified location (e.g., electronic devices and / or computing services communicatively coupled to the NICD system) (e.g., at least one network entity assigned to monitor and / or operatively control the specific network and / or network component, such as: an application server associated with an endpoint; a network management service; or a monitoring and / or notification system of an Internet service provider) and / or other electronic devices. The alert can indicate that the network component is believed to be the cause of the communication problem, the condition the network component is experiencing that is causing the communication problem, and / or possible mitigation actions to resolve the communication problem. The alert can be provided as a graphical user interface (GUI) alert, email, push notification, and / or other electronic communication.
[0016] In this way, even in highly complex networks with numerous intermediate components, the NICD system can quickly and accurately identify the location that may be the cause of network communication problems. This indication of the identified location can be automatically provided to the network entities associated with the identified location via alerts, thereby effectively notifying them of the possible causes of network communication problems.
[0017] In summary, Figure 1 This is a block diagram illustrating the components of a network-based communication system 100 that performs network problem causal analysis according to various aspects of this disclosure. The network-based communication system 100 may include various client devices 102 (e.g., electronic devices initiating network communication, such as personal computers (PCs), laptops, and / or servers), endpoints 104 (e.g., destination endpoints / electronic devices, such as PCs, laptops, and / or servers, serving as the target destination of network communication initiated by client devices 102), local networks of client devices 102 (e.g., local area networks (LANs) 106), remote networks (e.g., wide area networks (WANs) 108), and a network problem causal detection (NICD) system 110.
[0018] Client device 102 may attempt to communicate with endpoint 104 located remotely from client device 102. For example, client device 102 may attempt to communicate with endpoint 104 located on WAN 108, outside of client device 102's LAN 106.
[0019] Endpoint 104 can be any electronic device capable of interacting with client device 102 by sending and receiving electronic data. For example, endpoint 104 can be a server hosting applications such as Voice over IP (VoIP) or video conferencing applications. In other cases, endpoint 104 can be a user-facing device (e.g., another mobile phone, laptop computer). Endpoint 104 can communicate with multiple client devices 102. Although this description includes one endpoint 104, any number of endpoints 104 can be accessible by client devices 102 across LAN 106 and / or WAN 108.
[0020] Client device 102 can communicate with endpoint 104 by transmitting packets across LAN 106 and WAN 108 to endpoint 104. These packets may traverse through numerous intermediate routing devices (e.g., switches and / or routers) used to route the packets to endpoint 104. That is, LAN 106 and WAN 108 may consist of numerous intermediate routing devices used to facilitate communication between client device 102 and endpoint 104. Multiple client devices 102 may connect to LAN 106, and similarly, many client devices 102 may communicate via WAN 108. When a particular client device 102 encounters a communication problem (e.g., stagnation, crash), locating the root cause of the communication problem can be useful. For example, the location of the communication problem can be provided to the network service provider of LAN 106 and / or WAN 108 to troubleshoot the communication failure and restore communication to the desired operational level.
[0021] The NICD system 110 may be a computing device such as a server (e.g., having a processor 116). In some cases, the functionality of the NICD system 110 may be implemented as instructions implemented by an application and / or other computing device (e.g., server-based applications, cloud-based applications) stored on a non-transitory computer-readable medium 118. When executed by the processor 116, the computer-implemented instructions can implement the functionality of the NICD system 110 described herein. First, it should be noted that, although Figure 1 The system 100 depicted in the illustration shows a NICD system 110 communicatively connected to a LAN 106, but the NICD system 110 can also be implemented at any network level. For example, the NICD system 110 can be implemented at a WAN 108, a Personal Area Network (PAN), a Campus Network (CAN), a Metropolitan Area Network (MAN), or any other suitable network environment.
[0022] The NICD system 110 can facilitate causal analysis of this network communication problem to identify the location causing the problem. To perform this causal analysis, the NICD system 110 can receive Performance Problem Indication (IPI) data 112, which may be performance metric data indicating a performance problem originating from LAN 106. In some cases, such as in Transmission Control Protocol (TCP) packets and / or Precision Time Protocol (PTP) packets, the IPI data 112 can be extracted from the headers of these packets. For example, in Transmission Control Protocol (TCP) communication, the TCP packet header may include TCP IPI data 112. The NICD system 110 can use this IPI data 112 to determine the type of communication problem associated with the communication session between client device 102 and endpoint 104.
[0023] Continuing the description of network communication system 100, NICD system 110 can continuously receive data, such as IPI data 112 from LAN 106. For example, IPI data 112 can be received in the TCP packet header whenever packets are exchanged between client device 102 and endpoint 104 according to TCP communication. In some cases, NICD system 110 can request and / or receive IPI data 112 in response to reported communication problems associated with client devices.
[0024] The NICD system 110 can also periodically (e.g., at fixed intervals) receive data from intermediate routing devices. For example, the NICD system 110 may receive queue monitoring data 114 (e.g., queue status data indicating that packets are dropped or delayed at the ports of intermediate routing devices) incrementally from intermediate routing devices in LAN 106 and / or WAN 108. That is, the NICD system 110 may receive data every second, every ten seconds, every minute, and / or any other specified time range. In some cases, the NICD system 110 may request or retrieve data from LAN 106. For example, if the NICD system 110 identifies a communication problem based on IPI data 112, it can access queue monitoring data 114 of the set of intermediate routing devices associated with the communication session path between client device 102 and endpoint 104 to further pinpoint the cause of the communication problem. In other words, in some cases, based on IPI data 112 combined with queue monitoring data 114, the NICD system 110 can identify communication problems in a specific communication session between client device 102 and endpoint 104, or between a group of client devices 102 and endpoint 104. Furthermore, using IPI data 112 in conjunction with queue monitoring data 114, the NICD system 110 can identify the possible location causing the communication problem in that specific communication session (e.g., LAN 106 or WAN 108 and / or specific components of LAN 106 or WAN 108). Specifically, the possible location can be identified by observing data patterns in IPI data 112 and / or queue monitoring data 114, as detailed below.
[0025] Upon identifying the location of a communication problem and / or its cause, the NICD system 110 can generate and provide an alert 120. For example, the alert may include a graphical user interface (GUI) dialog box, a push notification to a specific electronic device (such as client device 102, endpoint 104, and / or monitoring services of system 100). In some cases, the alert may include an indication of the network problem and its identified location. In this way, notification of the cause of the network problem can be provided quickly and effectively for rapid response.
[0026] Now let's move on to a more detailed example. Figure 2 The illustration shows a client device 202 (e.g., Figure 1 Client device 102) and endpoint 204 (e.g., Figure 1 A block diagram of system 200 showing the detailed network path between endpoints 104 in LAN 206. For example, in the current example, client device 202 crosses LAN 206 (e.g., ...). Figure 1 LAN 206 (as described in the example) communicates with endpoint 204. LAN 206 may include various intermediate routing devices that form at least a portion of the network path between client device 202 and endpoint 204 (e.g., facilitating network communication). For example, LAN 206 may include intermediate routing devices such as access switch 208A, aggregation switch 208B, and core switch 208C on the network path between client device 202 and endpoint 204. In some cases, LAN 206 may include additional and / or alternative types of intermediate routing devices, such as WAN edge devices, on this network path. LAN 206 may also include intermediate routing devices 208D and 208E, which are not part of the network path between client device 202 and endpoint 204. When client device 202 wishes to communicate with endpoint 204, it may transmit packets destined for endpoint 204 (e.g., as indicated by destination information in the packet header) to access switch 208A of LAN 206. Access switch 208A can read the header from packets and transmit them to the appropriate destination, such as aggregation switch 208B. Each packet can be transmitted across LAN 206 until it reaches an intermediate routing device (e.g., core switch 208C) at the edge of LAN 206. The packet can then travel through other networks such as WAN 216 (e.g., ...). Figure 1 The packet is transmitted through WAN 108 to reach endpoint 204. After traversing WAN 216 and possibly other networks, the packet can be received by endpoint 204.
[0027] A communication session can include multidirectional communication. For example, client device 202 can send packets to endpoint 204, and endpoint 204 can send packets to client device 202. In some cases, this communication can occur on separate network paths. For example, dynamic routing and load balancing can affect the network path between client device 202 and endpoint 204. In this way, the network paths of packets sent by client device 202 and packets received by client device 202 can be different.
[0028] NICD system 210 (e.g., Figure 1 The NICD system 110 may include a non-transitory computer-readable medium (CRM) 218 (e.g., Figure 1The CRM 118 in the NICD system 210 stores computer-readable instructions on a non-transitory computer-readable medium. These computer-readable instructions are processed by the processor(s) 220 (e.g., ...) of the NICD system 210. Figure 1 When the (multiple) processors 116) execute, the NICD system 210 performs causal analysis of network communication problems. For example, the NICD system 210 can accumulate data from intermediate routing devices in LAN 206. For example, the NICD system 210 can receive TCP / IPI data 212 from access switch 208A (e.g., Figure 1 The TCP IPI data 212. In some cases, the NICD system 210 may receive Transmission Control Protocol (TCP) IPI data 212 from the access switch 208A. The TCP IPI data 212 may refer to various categories of data associated with a TCP-based communication session across the access switch 208A between client device 202 and endpoint 204. The TCP IPI data 212 may provide an indicator of bidirectional data flow (e.g., from client device 202 to endpoint 204, or from endpoint 204 to client device 202).
[0029] The NICD system 210 can also receive queue monitoring data 214 from various intermediate routing devices in the network path. Each intermediate routing device may have an Output Port Queue (OPQ). For example, access switch 208A may receive packets from client device 202 and perform a forward lookup to determine the appropriate OPQ to which the packet should go. Access switch 208A may add the packet to the OPQ with the appropriate priority using the Differentiated Service Code Point (DSCP) value of the packet header. The packet scheduler (not shown) may then dequeue the packet and pass it out of the OPQ (e.g., to be sent to another intermediate routing device (e.g., from access switch 208A to aggregation switch 208B). This process may be repeated by each intermediate routing device in the network path. That is, in order to organize the packets received by the intermediate routing devices, OPQs are used to determine the order in which the received packets are processed and transmitted to the target destination. Intermediate routing devices may have multiple OPQs associated with different communication sessions, different communication priorities, etc. For example, an intermediate routing device may have a high-priority OPQ for communication sessions that are considered high-priority.
[0030] NICD system 210 can receive queue monitoring data 214 for queues associated with network communication from intermediate routing devices in the network path between client 202 and endpoint 204. For example, in this system 200, the NICD system can receive queue monitoring data 214 from access switch 208A, aggregation switch 208B, and core switch 208C. NICD system 210 can retrieve queue monitoring data 214 for specific queues used by network communication for cause-and-effect analysis of communication problems. Although this example depicts three intermediate routing devices in the network path, queue monitoring data 214 can also be obtained from various other types of intermediate routing devices, such as additional switches, routers, WAN edge devices, and other devices in the network path. NICD system 210 can analyze queue monitoring data 214 in conjunction with TCP IPI data 212 to pinpoint the cause of communication problems.
[0031] In response to identifying the location of the communication problem, the NICD system 210 can provide an alert 222 to one or more network entities (e.g., Figure 1 Alert 120). For example, here, a client device 202 experiencing a network communication problem is provided with an alert 202, which can indicate the location of the cause of the network communication problem.
[0032] Figures 3 to 6 Procedures are provided for combining IPI data with queue monitoring data to pinpoint the location of network problem causes. These procedures can be implemented by NICD systems (e.g., Figure 1 NICD system 110 and / or Figure 2 This is implemented using the NICD system 210. For example, Figures 3 to 6 The process can be achieved through Figure 1 The computer-readable medium 118 stores and transmits Figure 1 The computer-implemented instructions and / or instructions executed by the (multiple) processors 116 Figure 2 The computer-readable medium 218 stores and transmits Figure 2 The computer implements instructions executed by (multiple) processors 220.
[0033] Figure 3 A flowchart is provided for process 300 for identifying and locating the cause of a network communication problem. Although the following description of process 300 is described as being performed by a NICD system (e.g., Figure 1 NCID system 110 Figure 2The process 300 described herein is executed by the NICD system 210, but it should be noted that any suitable device capable of receiving and processing data can execute the process 300 described herein. Furthermore, although the process 300 is described in a specific order, it should be understood that the process 300 can be executed in any suitable order, and one or more boxes described herein may be excluded.
[0034] As described in detail below, the NICD system can analyze TCP IPI data to determine conditions associated with client devices and endpoints. Therefore, at box 302, the NICD system receives a TCP IPI dataset (e.g., Figure 1 IPI data 112, Figure 2 TCP / IPI data 212), the dataset may include information about the client device (e.g., Figure 1 Client device 102, Figure 2 The client device 202) and the endpoint (e.g., Figure 1 Client device 102, Figure 2 The performance metrics data of network communication between endpoints 204. For example, the NICD system can obtain performance metrics data from the access switch (e.g., the one with which the client device is communicating) from the access switch. Figure 2 The access switch 208A receives TCP / IPI data. TCP / IPI data can be received in the header of TCP packets.
[0035] TCP IPI data can provide different information at different stages of communication between a client device and an endpoint. For example, a NICD system may receive a first portion of the TCP IPI dataset during the connection establishment / TCP handshake phase between a client device and an endpoint. A NICD system may also receive a second portion of the TCP IPI data, which may indicate a stable communication phase between the client device and the endpoint. The following paragraphs summarize the different categories of TCP IPI data (e.g., performance metrics). However, it should be noted that additional categories of TCP IPI data may also be considered in the causal analysis of the communication problems presented herein, and therefore, these categories of TCP IPI data also fall within the spirit of this disclosure.
[0036] Initial state IPI data
[0037] One category of IPI data can be the initial round-trip time (RTT) of the TCP handshake. The initial RTT indicates the time range used for a client device to connect to an endpoint. A TCP handshake begins with the client device sending a SYN packet to the endpoint. The endpoint can respond with a SYN ACK packet. The client device then responds with an ACK packet, thus completing the TCP handshake between the client device and the endpoint. The initial RTT can refer to the amount of time between sending the SYN packet to the endpoint and receiving the SYN ACK packet from the endpoint. For example, the initial RTT can be determined based on timestamps on Precision Time Protocol (PTP) packets or TCP packets.
[0038] Another category of IPI data can be the Initial Receive Window Size (RWS). RWS refers to the amount of data (e.g., bytes) that the client device and endpoint can accept. That is, the client device and endpoint can have buffers (e.g., temporary storage devices) that define the amount of data each can process within a given time. For example, a client device might have an initial buffer size available for receiving data (e.g., megabytes, kilobytes). The endpoint can dynamically adjust the RWS passed to the client device based on the dynamic amount of data it can receive. Conversely, the client device can also dynamically adjust the RWS passed to the endpoint. In some cases, intermediate routing devices along the network path of network communication (e.g., Figure 2 The access switches (208A, aggregation switches (208B), and core switches (208C) can also dynamically adjust the RWS passed to client devices and / or endpoints depending on the packet flow direction.
[0039] Another category of initial IPI data can be the Maximum Segment Size (MSS). TCP packets can consist of a header (e.g., an Internet Protocol (IP) header and a TCP header) and segments. A segment can be a unit of data carried by the TCP packet. The MSS is the size of the segments that an endpoint can send to a client device and that the client device can send to the endpoint. In some cases, intermediate routing devices in the network path can have different MSS parameters. If the segment size of a packet exceeds the MSS parameter of a particular intermediate routing device, the packet can be dropped or fragmented. Therefore, during the TCP handshake, the endpoint can adjust or limit the MSS that the client device can send to it. For example, if the MSS parameter of three intermediate routing devices is 1460 bytes and the MSS parameter of a fourth intermediate routing device is 1220 bytes, the endpoint can set the MSS to 1220 to prevent packet drop or fragmentation.
[0040] The initial category of IPI data can also be associated with the OPQ cache. As mentioned above, during the TCP handshake, when an access switch receives a packet from a client device, it can perform a forward lookup to determine the appropriate OPQ on the access switch to process the packet. This process can be repeated for each intermediate routing device in the network path. OPQ cache data refers to the identifier of a specific OPQ for network communication.
[0041] Steady-state IPI data
[0042] The NICD system can also receive and analyze IPI data throughout the entire communication session between the client device and the endpoint. An example of a steady-state IPI data category could be an ongoing RTT. As described regarding the initial RTT, the client device and endpoint communicate using acknowledgments. That is, whenever a packet is received, the receiver should send an acknowledgment to the sender. An ongoing RTT refers to the time interval between the sender sending a packet and the sender receiving an acknowledgment of that packet. In this way, an ongoing RTT can provide continuous tracking of the RTT during steady-state communication.
[0043] Another category of steady-state IPI data can be ongoing RWS. Similar to the initial RWS, ongoing RWS refers to the maximum buffer that a client device or endpoint can accept at any given time. The buffer on the client device or endpoint can change during a communication session. For example, if a client device initiates multiple applications after a TCP handshake, its buffer may decrease. Therefore, the client device can dynamically adjust its ongoing RWS. Consequently, an endpoint can reduce the number of packets it can send to the client device before receiving an acknowledgment.
[0044] The additional category of steady-state IPI data can be an indication of TCP re-xmit. As mentioned above, the TCP protocol specifies that acknowledgments should be made when sending packets. Acknowledgments can be based on the sequence number of the sent packet. The sequence number can be stored in the packet header to indicate the first byte of data in the packet. For example, by sending an ACK packet corresponding to the sequence number of a received packet, a client device can acknowledge that it has received the packet from the endpoint. If the sender does not receive an acknowledgment, it can retransmit the original packet with the same sequence number. That is, a TCP re-xmit can indicate that the OPQ of one or more intermediate routes is full (causing it to drop packets), and that the sender is now performing slow start (e.g., retransmitting TCP segments of dropped packets based on unacknowledged packet sequence numbers). A TCP re-xmit can refer to the total number of packets retransmitted during a communication session, the percentage of retransmitted packets (e.g., one percent of all packets sent during a communication session are retransmitted), or the rate at which packets are retransmitted during a communication session (e.g., ten packets retransmitted within a five-second interval).
[0045] Steady-state IPI data may also include the TCP Congestion Window Reduction (CWR) flag. The CWR flag is an indication provided to client devices that they should reduce their transmission rate. The CWR flag can be in response to Explicit Congestion Notification Echo (ECE) bits included in the TCP header of one or more packets. ECE bits can be generated by intermediate routing devices (such as access switches) along the network path to indicate that the intermediate routing device is experiencing network congestion. For example, this can occur when the OPQ on the network path between the client device and the endpoint is at or near full capacity. The CWR flag differs from other categories of steady-state IPI data (e.g., ongoing RTT and ongoing RWS) because it is associated with an intermediate routing device experiencing congestion, rather than with the endpoint.
[0046] Now returning to the description of process 300 for identifying and locating network communication problems, at box 304, the NICD system can receive queue monitoring data (e.g., from multiple intermediate routing devices along the network communication path between the client device and the endpoint) Figure 1 Queue monitoring data 114 Figure 2 (Queue monitoring data 214). That is, the NICD system can receive queue monitoring data about specific OPQs, across which client devices and endpoints communicate. In some cases, the NICD system can receive queue monitoring data periodically. In other cases, the NICD system can request queue monitoring data from a subset of intermediate routing devices based on the analysis of IPI data (e.g., when the analysis of IPI data indicates a network communication problem).
[0047] An example of queue monitoring data can include queue packet drop (Q-Drop). Q-Drop can refer to the total number of packets dropped or the rate at which packets are dropped within a corresponding OPQ at an intermediate routing device within a given time period. Q-Drop can also be analyzed at different levels of granularity. That is, the NICD system can receive Q-Drop data for a specific OPQ used by client devices and endpoints.
[0048] Another example of queue monitoring data can include queue utilization (Q-Util). Each intermediate routing device can monitor and report its OPQ utilization percentage for a defined time range. Q-Util refers to the percentage of an intermediate routing device's OPQ that is occupied during that interval. For example, an OPQ might be able to hold 100 packets in a given time period. If that OPQ holds an average of 50 packets during a defined interval (e.g., 5 seconds), then the Q-Util is 50 percent.
[0049] Another example of queue monitoring data can refer to queue delay (Q-Delay). Q-Delay can refer to the delay and variation in OPQ from ingress to egress. When packets arrive at intermediate routing devices, they may be separated into OPQs. The amount of time a packet spends in an OPQ can be delayed (e.g., in milliseconds or microseconds). Q-Delay can also refer to the high (e.g., above a defined threshold) average amount of time each packet spends in an OPQ (e.g., packets spend an average of one second in a particular OPQ). Q-Delay can also refer to queue jitter. Queue jitter is a measure of the variation or fluctuation in the amount of time a packet spends in an OPQ. That is, one packet may spend one millisecond in an OPQ, while a second packet may spend 500 milliseconds in the same OPQ. Q-Delay can be determined based on timestamped Precision Time Protocol (PTP) packets. For example, an intermediate routing device may periodically send loopback PTP packets to its own OPQ. When a PTP packet reaches the top of the queue (e.g., about to be sent), it is timestamped. Instead of being transmitted to another intermediate routing device, PTP packets are looped back for processing by the same intermediate routing device in which they are queued. PTP packets can be used to calculate queue time (e.g., the amount of time a PTP packet spends in the queue, calculated based on its timestamp) to determine if there is delay and / or fluctuation compared to previous measurements. In practice, Q-Delay can affect streaming or voice communications. For example, Q-Delay can be associated with output lag during a conference call, causing a presenter's voice to sound rapidly sped up.
[0050] At box 304, based on TCPIPI dataset 212 combined with queue monitoring data, the NICD system identifies specific components of network communication as the cause of communication problems between client devices and endpoints. For example, a specific component of the network could be one or more intermediate routing devices among intermediate routing devices on the network path. This specific component could also be an endpoint, such as a communication problem occurring on an application server. Alternatively, this specific component could be a service provider, such as an Internet service provider associated with the WAN on the path between the client device and the endpoint.
[0051] Trends or patterns in IPI and queue monitoring data can indicate specific components that may be causing communication problems. In cases where the NICD system automatically receives IPI and queue monitoring data, this analysis can be performed automatically. In other cases, the NICD system can analyze IPI and queue monitoring data when notified of a communication problem (e.g., by a client device or system administrator).
[0052] Turn to specific patterns in IPI data and queue monitoring data that indicate the cause of network communication problems in specific locations. Figures 4 to 6 Examples of using IPI data to determine specific components as the cause of communication problems are provided. While these examples involve specific categories of IPI data, queue monitoring data, and calculations, additional parameters and calculations used alone or in combination with queue monitoring data can also indicate communication problems. Figures 4 to 6 Each process in the process is illustrated as a separate process, but in some cases, all processes or subsets of processes can be combined to identify different locations where different network communication problems occur.
[0053] For example, Figure 4A flowchart is provided depicting a process 400 using Performance Problem Indicator (IPI) data to determine whether a communication problem has occurred at a communication session endpoint. At block 402, the NICD system identifies at least one of the following from the TCP IPI set: the ongoing RWS is zero, the ongoing RWS has fallen below a threshold, or the ongoing RWS has decreased proportionally to the initial RWS at a rate exceeding the threshold. The NICD system can identify the ongoing RWS as zero when the endpoint is not accepting any packet transmissions. For example, this can occur when the endpoint's buffer is full or overloaded. The NICD system can identify the ongoing RWS as falling below a threshold, which can be a predetermined minimum available storage on the endpoint buffer. For example, when the threshold for ongoing RWS is set to 5 kilobytes, the NICD system can identify the ongoing RWS at the endpoint as 1 kilobyte. Furthermore, the NICD system can identify the ongoing RWS as decreasing proportionally to the initial RWS at a rate exceeding the threshold (e.g., 10%, 20%). This can occur when the available storage on the endpoint buffer decreases during a communication session between the client device and the endpoint.
[0054] At box 404, in response to one of the conditions identified in box 402, the NICD system determines that the communication problem occurs at an endpoint of the path. That is, in response to identifying that the ongoing RWS is zero, the ongoing RWS has fallen below a threshold, and / or the ongoing RWS decreases proportionally to the initial RWS exceeding the threshold rate, the NICD system determines that the endpoint is a specific component of the network communication that is the cause of the communication problem between the client device and the endpoint. In some cases, the NICD system can make this determination without analyzing any queue monitoring data. In fact, at box 402, the NICD system identifies that the endpoint buffer (as determined by the ongoing RWS) is experiencing a state of being full or at least decreasing. For example, if the endpoint is an application server, this can indicate that the problem is occurring at that server, and consequently, that intermediate routing devices and other components of the network path are not the cause of the communication problem.
[0055] As mentioned above, the NICD system can also use queue monitoring data to identify specific components of network communication that are causing network congestion problems. Figure 5A sample flowchart 500 is provided depicting the process of using IPI data in conjunction with queue monitoring data. At box 502, the NICD system identifies a CWR flag or TCP Re-xmit indication at the access switch from the TCPIPI dataset. As described above, a CWR flag may appear when network congestion occurs at one or more intermediate routing devices in the path between the client device and the endpoint. The CWR flag instructs the client device to slow its transmission rate to reduce the load on the OPQ of at least one of the intermediate routing devices. At this box, the NICD system may also identify whether the total number of packets being retransmitted or the packet transmission rate exceeds a certain threshold. That is, the NICD system can identify that TCP Re-xmit is exceeding a threshold.
[0056] At box 504, in response to one of the conditions identifying box 502, the NICD system determines that the communication problem is associated with network congestion. For example, as mentioned above, the CWR flag can instruct the client device to reduce its transmission rate. This request can indicate that the network is unable to handle the client device's previous transmission rate, and therefore network congestion exists. Furthermore, packet retransmission can be triggered based on dropped packets, indicating that the network is unable to handle all packets provided by the client device, and therefore network congestion exists.
[0057] Proceeding to decision box 506, the NICD system queries queue monitoring data to identify whether a Q-Drop exists at any intermediate routing device on the network path between the client device and the endpoint during a common time window of network congestion (e.g., a predefined surrounding time range). The NICD system can receive queue monitoring data from intermediate routing devices on the network path between the client device and the endpoint. After IPI data indicates the presence of a network congestion problem, the NICD system can query the monitoring data to determine whether any Q-Drop and / or Q-Util exceeds a predefined threshold on the relevant OPQ of the intermediate routing device on the network path between the client device and the endpoint. In some cases, the NICD system can analyze the queue monitoring data for breaches of a threshold number of Q-Drops and / or Q-Utils at each OPQ used for network communication (e.g., a communication session) during the relevant common time window of network congestion.
[0058] To perform analysis efficiently and effectively, the NICD system can analyze a subset of queue monitoring data occurring around a threshold time period (e.g., a common time window) surrounding identified network congestion. For example, if the NICD system receives queue monitoring data periodically, it can analyze queue monitoring data received in the two periods before and after the NICD system identifies network congestion. Alternatively, the NICD system can query queue monitoring data for a threshold amount of time (e.g., one minute) before and / or after the identified network congestion time.
[0059] If Q-Drop is not present at any of the intermediate routing devices during the common time window of network congestion, process 500 continues at box 508, where the NICD system determines that the specific component causing the communication problem is located outside the LAN (e.g., at the network's service provider). For example, the specific component causing the communication problem could be identified as the Internet service provider of the WAN through which packets are transmitted.
[0060] Conversely, if a Q-Drop occurs at a specific intermediate routing device during the common time window of network congestion, process 500 proceeds to box 510. At box 510, the NICD system determines that the intermediate routing device experiencing the Q-Drop is a specific component causing the communication problem in the network. In some cases, multiple intermediate routing devices may be experiencing Q-Drops during the period of network congestion. In this case, the NICD system can determine that the multiple intermediate routing devices experiencing Q-Drops are specific components causing the communication problem in the network.
[0061] As mentioned above, regarding Figure 5 The process described is an analysis workflow used to determine which specific component in network communication is causing a communication problem. Figure 6 Another example flowchart is provided, depicting a process 600 that combines IPI data with queue monitoring data to determine the location of network latency communication problems.
[0062] At box 602, the NICD system identifies rate changes in the ongoing RTT within the TCP IPI set. As described above, the ongoing RTT is the amount of time elapsed from when the sender (e.g., a client device or endpoint) transmits a packet to when the sender receives an acknowledgment. The NICD system can identify when changes in the ongoing RTT exceed a certain threshold (e.g., seconds, milliseconds). This rate change can indicate faster communication (e.g., lower ongoing RTT) or slower communication (e.g., higher RTT).
[0063] Proceeding to box 604, in response to the conditions identified in box 602, the NICD system determines that the communication problem is associated with network latency. That is, rate variations in the ongoing RTT may be associated with network latency caused by intermediate routing devices in the path between the client device and the endpoint.
[0064] At decision box 606, the NICD system queries queue monitoring data to identify whether, during the common time window of network latency, there is a Q-Delay (e.g., an unwanted queuing phase from ingress to egress in an OPQ for network communication packets) and / or rate variation (e.g., the rate at which the queuing phase changes over time exceeds a predefined threshold) in any intermediate routing device's OPQ. As described with respect to decision box 506, the NICD system can examine queue monitoring data occurring before and / or after a threshold time before the identified network latency time. Using this queue monitoring data, the NICD system can determine whether a Q-delay exists at any OPQ in the relevant OPQ of the intermediate routing device between the client device and the endpoint within the common time window of network latency. For example, if an OPQ is experiencing slow output exceeding a certain threshold (e.g., 250 milliseconds), or if an OPQ is experiencing jitter, such as OPQ processing phase fluctuations (e.g., packet queuing time ranging from 1000 milliseconds to 400 milliseconds), the NICD system can determine that the OPQ is experiencing a Q-Delay.
[0065] If no Q-Delay or rate change is found in the queue output at the intermediate routing device during the common time window of network latency, process 600 proceeds to box 608. At box 608, the NICD system determines that the service provider associated with the network path is the specific component in the network communication that is causing the communication problem.
[0066] If a Q-Delay and / or rate change exists at an intermediate routing device during the common time window of the identified network delay, process 600 proceeds to box 610. At box 610, the NICD system determines that the intermediate routing device experiencing a Q-Delay or rate change is a specific component causing the communication problem in the network. Similar to box 510 above, in some cases, multiple intermediate routing devices may be experiencing a Q-Delay during the common time period of the network delay. In this case, the NICD system can determine that multiple intermediate routing devices experiencing a Q-Delay are specific components causing the communication problem in the network.
[0067] Reference Back Figure 3Now, turning to box 306, the NICD system can generate and provide an alert that indicates a specific component identified as the cause of a communication problem. This alert can be sent to a wide variety of entities. For example, if the specific component causing the network communication problem is an endpoint, both the client device and the endpoint can be notified. Therefore, if the endpoint is an application server, the entity running the application or hosting the application server can be notified. If the specific component causing the communication problem is an intermediate routing device, the local system administrator or network management service can be notified. Furthermore, if the specific component causing the communication problem is the WAN service provider, that service provider can be notified.
[0068] Notifications can indicate how communication problems are identified, their possible causes, and suggested mitigation measures. These examples illustrate the benefits of the techniques described herein, which allow entities associated with communication problems affecting a communication session to be notified of the issue. This facilitates faster network troubleshooting and communication problem resolution by enabling responsible entities to address any conditions causing the problem before it escalates or worsens (e.g., leads to a complete communication failure). Furthermore, it provides a positive indication to other entities associated with the network communication session that their network components are innocent and have not experienced any conditions that negatively impact the client.
[0069] It is understood that the current technology offers significant value. Specifically, it provides a network problem causality detection system that identifies potential locations between client devices and endpoints that may be causing network communication problems. Furthermore, even in highly complex networks with numerous intermediate components, the current technology can quickly and accurately identify the locations that may be the causes of network communication problems.
[0070] While certain features of this disclosure have been described and illustrated herein, many modifications and alterations will occur to those skilled in the art. Therefore, it should be understood that the appended claims are intended to cover all such modifications and alterations that fall within the true spirit of this disclosure.
Claims
1. A non-transitory computer-readable medium comprising computer-readable instructions that, when executed by one or more processors of one or more computers, cause the one or more computers to: Receive Transmission Control Protocol Performance Issue Indicator (TCP IPI) dataset, wherein the TCP IPI dataset includes performance metrics associated with network communication between client devices and endpoints; Receive queue monitoring data from multiple intermediate routing devices along the path of the network communication between the client device and the endpoint; Based on the TCPIPI dataset and the queue monitoring data, specific components on the path of the network communication are identified as the cause of communication problems between the client device and the endpoint. as well as An alert is generated and provided, which indicates the specific component that is the cause of the communication problem.
2. The non-transitory computer-readable medium of claim 1, wherein the non-transitory computer-readable medium comprises computer-readable instructions, which, when executed by the one or more processors of the one or more computers, cause the one or more computers to: The specific component is selected from a set of network components for the network communication, wherein the set of network components includes: The plurality of intermediate routing devices on the path The endpoints of the path, and The service provider associated with the path.
3. The non-transitory computer-readable medium of claim 1, wherein the performance measurement data includes at least one of the following: Round-trip time (RTT) indicated by one or more TCP packets; The receive window size RWS of the one or more TCP packets, The output port queue of the access switch for the path is used to cache data. Indicator of TCP packets being retransmitted; The maximum segment size (MSS) indicated by the one or more TCP packets; or The CWR flag is reduced by the congestion window indicated by the one or more TCP packets.
4. The non-transitory computer-readable medium of claim 3, wherein the non-transitory computer-readable medium comprises computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: During the connection establishment phase between the client and the endpoint, a first portion of the TCP IPI dataset is received, wherein the first portion of the TCP IPI dataset includes at least one of the following: The initial RTT is indicated by the initial set of the one or more TCP packets transmitted during the connection establishment phase; The initial RWS, indicated by the initial set of the one or more TCP packets; The output port queue caches data; or The MSS is indicated by the initial set of the one or more TCP packets; as well as During a stable-state communication phase between the client and the endpoint, a second portion of the TCP IPI dataset is received, wherein the second portion of the TCP IPI dataset includes at least one of the following: An ongoing RTT, indicated by a second set of the one or more TCP packets transmitted during the steady-state communication phase; An ongoing RWS, indicated by the second set of the one or more TCP packets; The indication that the TCP packet is being retransmitted to the access switch; or The CWR flag, indicated by the second set of the one or more TCP packets.
5. The non-transitory computer-readable medium of claim 4, wherein the non-transitory computer-readable medium includes computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: One or more conditions are identified from the TCPIPI set, the one or more conditions including at least one of the following: The ongoing RWS is zero; The ongoing RWS drops below the threshold; or The ongoing RWS decreases proportionally to the initial RWS at a rate exceeding a threshold; and In response to identifying the one or more conditions, it is determined that the endpoint of the path is the specific component in the network communication, and the specific component is the cause of the communication problem.
6. The non-transitory computer-readable medium of claim 4, wherein the non-transitory computer-readable medium comprises computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: One or more conditions are identified from the TCP IPI dataset, the one or more conditions including at least one of the following: the TCP IPI dataset includes the CWR flag, or the indication that TCP packets are being retransmitted to the access switch; and In response to identifying one or more of the conditions, it is determined that the communication problem is associated with network congestion.
7. The non-transitory computer-readable medium of claim 6, wherein the non-transitory computer-readable medium includes computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: In response to identifying the communication problem as being associated with the network congestion: Query the queue monitoring data to identify whether, during the common time window of the network congestion, there is a queue packet drop (Q-Drop) in any queue associated with the output port queue cache data at any of the plurality of intermediate routing devices; When Q-Drop exists at a specific intermediate routing device among the plurality of intermediate routing devices, the specific intermediate routing device is determined to be the specific component in the network communication, and the specific component is the cause of the communication problem; and When Q-Drop is not present at any of the plurality of intermediate routing devices, the service provider associated with the path is determined to be the specific component in the network communication that is the cause of the communication problem.
8. The non-transitory computer-readable medium of claim 4, wherein the non-transitory computer-readable medium comprises computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: The TCPIPI data identifies variations in the rate of the ongoing RTT; and In response to identifying the change in the rate of the TCPIPI data, which includes the ongoing RTT, the communication problem is determined to be associated with network latency.
9. The non-transitory computer-readable medium of claim 8, wherein the non-transitory computer-readable medium includes computer-readable instructions that, when executed by the one or more processors of the one or more computers, cause the one or more computers to: In response to determining that the communication problem is associated with the network latency: Query the queue monitoring data to identify whether, during the common time window of the network latency, there is at least one of latency or rate variation in the queue output at any of the plurality of intermediate routing devices; When there is at least one of delay or rate variation in the queue output at a specific intermediate routing device of the plurality of intermediate routing devices, the specific intermediate routing device is determined to be the specific component in the network communication, and the specific component is the cause of the communication problem. as well as When at least one of latency or rate variation is not present in the queue output at any of the plurality of intermediate routing devices, the service provider associated with the path is determined to be the specific component in the network communication that is the cause of the communication problem.
10. The non-transitory computer-readable medium of claim 1, wherein the TCPIPI data is extracted from at least one of: Transmission Control Protocol (TCP) packets or Precision Time Protocol (PTP) packets.
11. A processor-implemented method, comprising: Receive Transmission Control Protocol Performance Issue Indicator (TCP IPI) dataset, wherein the TCP IPI dataset includes performance metrics associated with network communication between client devices and endpoints; Receive queue monitoring data from multiple intermediate routing devices along the path of the network communication between the client device and the endpoint; Based on the TCPIPI dataset and the queue monitoring data, specific components on the path of the network communication are identified as the cause of communication problems between the client device and the endpoint. as well as An alert is generated and provided, which indicates the specific component that is the cause of the communication problem.
12. The processor-implemented method according to claim 11, comprising: The specific component is identified from a set of network components of the network communication, wherein the set of network components includes: A specific intermediate routing device among the plurality of intermediate routing devices on the path. The endpoints of the path, and The service provider associated with the path.
13. The processor-implemented method according to claim 12, comprising: In response to identifying the specific intermediate routing device as the specific component, the OPQ in the multiple output port queues OPQ on the specific intermediate routing device is determined as the cause of the communication problem.
14. The processor-implemented method of claim 12, wherein the particular component includes at least two of the plurality of intermediate routing devices.
15. The processor-implemented method of claim 11, wherein the alarm is provided to at least one network entity having operational control over the particular component, the network entity comprising at least one of the following: The application server associated with the endpoint; Network management services, or Internet service providers.
16. The processor-implemented method according to claim 11, comprising: From the TCPIPI dataset, the cause of the communication problem is determined to be associated with at least one of the following: network congestion or network delay.
17. The processor-implemented method of claim 16, wherein the alarm includes an indication that the cause of the communication problem is associated with at least one of: network congestion or network delay.
18. A Network Problem Causality Detection (NICD) system, comprising: processor; as well as A non-transitory computer-readable medium, the non-transitory computer-readable medium comprising computer-readable instructions, which, when executed by the processor, cause the processor to: Receive Transmission Control Protocol Performance Issue Indicator TCPIPI dataset, wherein the TCPIPI dataset includes performance metric data associated with network communication between client devices and endpoints; Receive queue monitoring data from multiple intermediate routing devices along the network path of the network communication; Based on the TCPIPI dataset and the queue monitoring data, specific components on the network path of the network communication are identified as the cause of communication problems between the client device and the endpoint. as well as An alert is generated and provided, which indicates the specific component that is the cause of the communication problem.
19. The system of claim 18, wherein the NICD system is configured to: Receive notification of the communication problem from the client device; and In response to receiving the notification of the communication problem from the client device, a request is made for the queue monitoring data from the plurality of intermediate routing devices.
20. The system of claim 18, wherein the NICD system is configured to: The TCPIPI dataset is extracted from TCP packets transmitted across at least a portion of the network path between the client device and the endpoint.