Packet sampling control for on-box anomaly detection and reporting and out-of-band message generation and transmission for use in root cause determination of latency service level agreement (SLA) violations
By generating messages that do not modify packets at communication network nodes, and performing non-random sampling and anomaly detection, the accuracy problem of root cause analysis of latency in large networks is solved, and the relevance and analysis precision of the sampled information are improved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JUNIPER NETWORKS INC
- Filing Date
- 2025-10-30
- Publication Date
- 2026-07-21
AI Technical Summary
In large-scale communication networks, it is difficult to determine the root cause of high latency. Existing technical methods suffer from sampling randomness and information processing challenges, leading to inaccurate latency violation analysis.
By generating messages in communication network nodes that do not modify packets, including absolute delay, expected delay, and cause of anomalies, for non-random sampling and efficient anomaly detection, and combining queue priority and congestion notification, control plane messages separate from sampled packets are generated.
It achieves efficient and accurate root cause analysis of latency, reduces the impact on packet bandwidth, and improves the correlation and analysis accuracy of sampling information.
Smart Images

Figure CN122437781A_ABST
Abstract
Description
Technical Field
[0001] This application relates to communication networks. In particular, this application relates to helping determine multiple root causes of data delays in at least a portion of a communication network. Background Technology
[0002] The discussion in this chapter is not intended to acknowledge existing technology. § 1.2.1 Latency Issues in Data Forwarding
[0003] When traffic used for applications / services experiences high latency (e.g., above threshold levels, such as those defined by the service agreement), there can be many potential causes or root causes for the high latency. For example, the range of problems can vary from link latency variations to queuing delays, application-level issues, and so on. It is crucial for network owners or operators to identify the problems(s) causing or contributing to high latency so they can take appropriate measures to mitigate or eliminate them. Unfortunately, however, identifying the causes(s) or root causes of latency variations in large networks is challenging.
[0004] To identify the multiple issues causing or causing high latency in a network, various proposed operational methods exist, depending on the actual data traffic sampled. These methods typically involve capturing the actual time spent on data traffic (e.g., packets) at each node and on each link to pinpoint the exact location where high latency occurs or has occurred. Unfortunately, these methods are often probabilistic and therefore may not truly capture latency data for the multiple packets suffering from latency violations. § 1.2.2 Existing Delay Detection Schemes and Their Perceived Limitations
[0005] Competitive solutions for detecting latency and / or determining the root causes of latency include: - J. Kumar et al., "Inband Flow Analyzer" draft-kumar-ippm- ifa-07 (Internet Engineering Task Force, September 7, 2023) (referred to as the “IFA Draft” and incorporated herein by reference); - H. Song et al., “Network Telemetry Framework” Draft for comments: 9232 (Internet Engineering Task Force, May 2022) (referred to as “RFC 9232” or “Network Telemetry Architecture”, and incorporated herein by reference); - C. Filsfils et al., "Path Tracing in SRv6 networks", draft-filsfils-spring-path-tracing-05 (Internet Engineering Task Force, October 23, 2023) (referred to as the “Path Tracing Draft” and incorporated herein by reference); and - H. Song et al., "On-Path Telemetry using Packet Marking to TriggerDedicated OAM Packets", draft-song- ippm-postcard-based-telemetry-16 (Internet Engineering Task Force, May 30, 2023) (Referred to as the “Draft Grouping Tags”, and incorporated herein by reference) Schemes such as those listed above are discussed below in §§ 1.2.2.1 and 1.2.2.2. § 1.2.2.1 In-band monitoring delay detection scheme
[0006] In-band monitoring methods have been proposed, such as: - F. Brockners, Ed., “In Situ Operations, Administration, and Maintenance (IOAM) Deployment” Draft for comments: 9378 (Internet Engineering Task Force, April 2023) (referred to as "RFC 9378", and incorporated herein by reference); - L. Andersson et al., "MPLS Network Actions Framework" draft-andersson-mpls-mna-fwk-01 (Internet Engineering Task Force, April 27, 2022) (referred to as the “MNA Draft” and incorporated herein by reference); and - “P4 Drives the Innovation and Practical Application of GlobalScheduling Ethernet GSE” (available online at / p4.org / p4-spec / docs / INT_v2_1.pdf, December 2, 2024) (also known as “INT-MD: (In-Band Telemetry Specification)”, and incorporated herein by reference.)
[0007] The method in the proposal cited above involves collecting operational and telemetry information from data packets as they traverse the path between two points in a communication network. Each of these monitoring methods defines a new additional header to record the dwell time. Typically, due to the high volume of traffic in the network, the collector cannot handle the dwell time of each packet. Since it is impossible to capture this information for all packets, flow selection and packet sampling mechanisms are used.
[0008] Unfortunately, these in-band monitoring schemes have many limitations. For example, they add information to the data packets (e.g., in the header). This added overhead increases the bandwidth of the stream and can therefore potentially exacerbate latency issues. Furthermore, these in-band monitoring schemes require new packet header processing. Therefore, they are typically not deployable until all nodes are able to process the new packet headers, which may require a hardware refresh of all nodes. § 1.2.2.2 Delay Detection Scheme for Exporting Telemetry Information from Nodes
[0009] Methods for deriving telemetry information from nodes have been proposed, such as: - H. Song et al., “On-Path Telemetry using Packet Marking to TriggerDedicated OAM Packets,” draft-song-ippm-postcard-based-telemetry-16 (Internet Engineering Task Force, May 30, 2023) (referred to as the “Packet Marking Draft” and incorporated herein by reference); and - “P4 Drives the Innovation and Practical Application of GlobalScheduling Ethernet GSE” (available online at / p4.org / p4-spec / docs / INT_v2_1.pdf, December 2, 2024) (also known as “INT-MD: (In-Band Telemetry Specification)”, and incorporated herein by reference.) The methods in these proposals collect metadata about (multiple) packets on the router, such as input / output interface, input / output time, etc. The collected information is then exported to an out-of-band external entity for analysis.
[0010] Deriving telemetry information from nodes typically presents numerous challenges. For example, using packet sampling mechanisms makes these solutions scalable. Unfortunately, the randomness or non-centralization of sampling makes it challenging to identify (and sample) the exact same packets on every box across the monitored communication path. Due to the use of packet sampling, it may be impossible to capture delay data for packets experiencing delay at each node. Therefore, root cause analysis of delay violations is often heuristic-based and may be inaccurate.
[0011] There are proposals to use certain bits in the packet header or a new header that carries an indication that the packet needs to be sampled, but these proposals themselves also bring similar problems and challenges to those discussed above when modifying the packet header to carry information.
[0012] Figure 1 A simple network topology 100 is used to illustrate a latency detection scheme for deriving telemetry information from nodes. As shown, customer edge devices CE1 120b and CE2 120b communicate via transport network 130. Transport network 130 includes provider edge devices PE1 140a and PE2 140b and relay routers 150a to 150c. The latency detection scheme includes at least one centralized collector 110 and agents 160a to 160e on each network node (e.g., router and / or switch) 140a, 140b, 150a, 150b, 150c in transport network 130. As shown by the dashed connector, agents 160a to 160e provide information to collector 110. Collectors 110 store some or all of the collected information and / or information derived from the collected information in a flow database 115.
[0013] The main architectural backbone, defined by nodes (PEs and transport routers) and (multiple) centralized collectors 110, implements network-wide, long-term, and / or short-term activity flow analysis. For example, it utilizes existing protocols (e.g., B. Claise, Ed., “Specification of the IP Flow Information Export (IPFIX) Protocol for the Exchange of Flow Information”). Draft for comments: 7011 (Internet Engineering Task Force, September 2013) (referred to as “RFC 7011”, and incorporated herein by reference) and other protocols discussed above.
[0014] As mentioned above, existing packet sampling settings are configured on each node in the cluster to sample (multiple) packets. Two characteristics of conventional sampling implemented in most existing routers / switches are: (1) a pre-configured sampling frequency on each node; and (2) each node performs sampling independently of the sampling performed by other nodes. The fact that a sampling frequency is pre-configured on each node makes it possible to randomly select packets from the ongoing flow for sampling. Due to the actual traffic density, sampling all frames is not feasible. The fact that each node performs sampling independently means that the (multiple) packets sent to the collector by different nodes, their characteristics, and / or information about them may differ. As mentioned above, this leads to independent flow observations, which may or may not be useful for detecting anomalies at the collector.
[0015] refer to Figure 2 Each node 210 (recall, for example, PE 140 and relay router 150) includes two main components in implementing flow monitoring: (1) a software application block 220; and (2) an ASIC packet pipeline 230. The software application block 220 may be part of the router's control plane (e.g., routing engine). The ASIC packet pipeline 230 may be part of the router's forwarding plane. In the software application block 220, software components in the agent (recall, 160) are involved in communicating with and configuring node 210 for sampling settings. The ASIC packet pipeline (e.g., packet parsing pipeline) 230 includes a sampling module / circuit system 232, which, as configured, performs actual sampling and provides the sampled information to the collector communication module 224 in the software block 220. As indicated by the dashed lines, the sampled information may be provided via a host path as complete packets, packet characteristics, and / or information about the packets. The stream sampled in the ASIC grouping pipeline 230 is currently random or decentralized, but the sampling rate set in the sampling configuration module 222 is confirmed.
[0016] refer to Figure 2 ASIC pipeline 230, Figure 3 The illustration shows an example packet forwarding engine (PFE) 300 that can be used to sample packets. The example PFE 300 includes an ingress interface 310, an ingress pipeline 320, a switching structure 330, an egress pipeline 340, and an egress interface 350. The PFE processes 322 / 344 of the ingress and / or egress pipelines 320 / 340 may include, for example, an ingress resolver, filters, samples, an ingress traffic manager, forwarding lookup, egress packet modification, an egress resolver, and an ingress result processor.
[0017] refer to Figure 3This section explains a typical flow monitoring implementation in a general-purpose ASIC. Packet sampling occurs in the ingress pipeline 320 and / or the egress pipeline 340 via flow monitoring traps 324 and / or 346, respectively. The ingress pipeline 320 includes a packet forwarding engine (PFE) process 322. The forwarding lookup process 322... i After completion, the received packets are queued in VOQ 326. Based on port credit, the packets are then dequeued and scheduled for egress processing. Finally, the packets leave the network port at egress interface 350 (e.g., via switching structure 330 and egress pipeline 340). Summary of the Invention
[0018] The example methods consistent with this specification can be used to help perform root cause analysis of traffic latency, or at least to intelligently collect information used in root cause analysis of traffic latency (e.g., in a targeted manner). An example method for a communication network node includes: (a) receiving ingress packets by the communication network node; (b) queuing the ingress packets received in a Virtual Output Queue (VOQ) by the packet forwarding engine (PFE) of the communication network node; (c) generating a delay measurement or indication by the communication network node based on at least one of: (A) queuing delay of the received ingress packets and / or (B) dwell time of the received ingress packets on at least a portion of the communication network node; (d) determining, by the communication network node, whether to sample the ingress packets using at least the generated delay measurement or indication; (e) generating a message (e.g., not within a packet and neither appended to nor preceding a packet) by the communication network node for each of at least one sampled ingress packet, the message including at least one of: (i) absolute delay, (ii) expected delay associated with the communication network node, and / or (iii) cause of an anomaly; and (f) sending the generated message toward a traffic sample collector by the communication network node for (e.g., root cause) analysis.
[0019] In at least some example implementations of the example methods, the dwell time of received inbound packets at at least a portion of the communication network node is normalized as a function of the expected delay associated with the communication network node. For example, the expected delay associated with the communication network node may be a function of at least one of the following: (A) the packet processing capability of the application-specific integrated circuit (ASIC) to which the inbound PFE belongs, (B) head-of-queue congestion at the queue, (C) the presence or absence of pass-through switching technology, (D) packet retrieval processing, (E) the type of on-chip memory used by the ASIC, (F) the type of off-chip memory used by the ASIC, (G) the port speed at which packets arrive at the port on the communication network node and / or (H) the port speed at which packets on the communication network node are forwarded to the port of the next hop.
[0020] In at least some example implementations of the example methods, the generated messages conform to RFC 7011.
[0021] In at least some example implementations of the example method, the generated message is separate from the sampled packet, and the action of sending the generated message toward the collector is separate from the action of the communication network node forwarding the sampled packet to the next hop.
[0022] In at least some example implementations of the example methods, the action of whether to sample ingress packets is determined by the communication network node at least using the generated delay measurement or indication, or based on packet priority or service level (CoS).
[0023] In at least some example implementations of the example method, the cause of the anomaly is at least one of the following: (A) queuing delay at the VOQ, (B) memory access delay on the entry PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline through the entry PFE of the communication network node. In at least some of these example implementations, the entry PFE is part of an application-specific integrated circuit (ASIC) configured to monitor and provide at least one of the following: (A) queuing delay, (B) memory access delay on the entry PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline through the entry PFE of the communication network node.
[0024] In some example implementations of the example method, the actions of queuing incoming packets received in the Virtual Output Queue (VOQ), generating delay measurements or indications, and determining whether to sample the incoming packets are performed by the PFE. In at least some of these example implementations, the PFE is an application-specific integrated circuit (ASIC) or a part thereof.
[0025] As mentioned above, the example methods consistent with this specification can be used to assist in performing root cause analysis of traffic latency, or at least to intelligently collect information used in root cause analysis of traffic latency (e.g., in a targeted manner). An example method for use on a collector device includes: (a) receiving, by the collector device, a plurality of flow congestion messages from a plurality of nodes in a communication network, wherein each of the plurality of flow congestion messages includes at least one of the following: (A) absolute latency experienced by the flow at the node, (B) expected latency for the node, and / or (C) an anomalous cause, and wherein each of the plurality of flow congestion messages is separate from the data packets of the flow; and (b) performing root cause analysis on the latency of one or more flows in the flow using the received plurality of IP flow messages. In at least some of these example methods, the samples included in the plurality of flow congestion messages are generated non-randomly by each of the plurality of nodes.
[0026] Example communication network nodes can be used to assist in performing root cause analysis of traffic latency, or at least intelligently collect information used in root cause analysis of traffic latency (e.g., in a targeted manner). An example communication network node includes: (a) an interface adapted to receive ingress packets; (b) a packet forwarding engine (PFE) adapted to (1) queue ingress packets received in a virtual output queue (VOQ), and (2) generate a latency measurement or indication based on at least one of: (A) queuing delay of ingress packets received by the communication network node and / or (B) dwell time of ingress packets received by the communication network node at least partially on the communication network node; and (c) an application-specific integrated circuit (ASIC) adapted to at least use the generated latency measurement or indication to determine whether to sample ingress packets.
[0027] In at least some example implementations, the example communication network node further includes: (d) at least one processor configured to (1) generate a message for each of at least one sampled ingress packet, the message including at least one of the following: (i) absolute delay, (ii) expected delay associated with the communication network node, and / or (iii) cause of anomaly; and (2) send the generated message toward a traffic sample collector for root cause analysis. In at least some of these example implementations of the communication network node, the cause of anomaly is at least one of the following: (A) queuing delay at the VOQ, (B) memory access delay on the ingress PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline of the ingress PFE of the communication network node.
[0028] In some example implementations of the example communication network, the dwell time of received inbound packets at at least a portion of the communication network node is normalized as a function of the expected delay associated with the communication network node. In at least some of these example implementations, the expected delay associated with the communication network node is a function of at least one of the following: (A) the packet processing capability of the application-specific integrated circuit (ASIC) to which the inbound PFE belongs, (B) head-of-queue congestion at the queue, (C) the presence or absence of pass-through switching technology, (D) packet retrieval processing, (E) the type of on-chip memory used by the ASIC, (F) the type of off-chip memory used by the ASIC, (G) the port speed at which packets arrive at the port on the communication network node and / or (H) the port speed at which packets on the communication network node are forwarded to the port of the next hop.
[0029] In some example implementations of the example communication network node, the generated messages conform to RFC 7011.
[0030] In at least some example implementations of the example communication network node, the generated message is separate from the sampled packet, and the action of sending the generated message toward the collector is separate from the action of the communication network node forwarding the sampled packet to the next hop.
[0031] In at least some example implementations of example communication network nodes, when at least the generated latency measurement or indication is used to determine whether to sample inbound packets, the ASIC is suitable for also considering packet priority (e.g., based on service level, queue priority, etc.).
[0032] A non-transitory computer-readable medium may be provided with processor-executable instructions that, when executed by at least one processor, perform any of the described methods. Attached Figure Description
[0033] Figure 1 It is a simple network topology used to illustrate a delay detection scheme for deriving telemetry information from nodes.
[0034] Figure 2 This is a block diagram of nodes with software application blocks and ASIC grouping pipelines used in the implementation of flow monitoring.
[0035] Figure 3 The illustration shows a typical flow monitoring implementation in a general-purpose ASIC.
[0036] Figure 4 An example network topology that can be deployed in accordance with the example embodiments described herein is illustrated.
[0037] Figure 5 This is a flowchart of an example method for performing packet processing and monitoring at a communication network (forwarding) node (e.g., a switch or router).
[0038] Figure 6 This is a flowchart of an example method for performing grouping analysis at the collector in a manner consistent with this specification.
[0039] Figure 7 The illustration shows two data forwarding systems that can be used as nodes coupled via communication links in a communication network (such as a communication network employing one or more traffic monitoring features of this application).
[0040] Figure 8 It is a block diagram of a router that can be used in communication networks (such as communication networks employing one or more traffic monitoring features of this application).
[0041] Figure 9This is a block diagram of an exemplary machine that can execute one or more of the processes described and / or store information used and / or generated by such processes. Detailed Implementation
[0042] This disclosure may relate to novel methods, apparatus, message formats, and / or data structures to aid in performing root cause analysis of traffic latency, or at least intelligently acquiring information used in root cause analysis of traffic latency (e.g., in a targeted manner). The following description is presented to enable those skilled in the art to make and use the described embodiments and is provided in the context of a particular application and its requirements. Therefore, the following description of exemplary embodiments provides illustration and description but is not intended to be exhaustive or to limit this disclosure to the precise forms disclosed. Various modifications to the disclosed embodiments will be apparent to those skilled in the art, and the general principles set forth below can be applied to other embodiments and applications. For example, although a series of actions may be described with reference to a flowchart, the order of actions may differ in other implementations when the execution of one action is independent of the completion of another action. Furthermore, independent actions may be performed in parallel. Unless explicitly described as such, elements, actions, or instructions used herein should not be construed as critical or essential to this specification. Moreover, as used herein, the article “a” is intended to include one or more items. Where only one item is intended, the term “a” or similar language is used. Therefore, this disclosure is not intended to be limited to the embodiments shown, and the inventors consider their invention to be the subject of any patentable subject described.
[0043] The example methods and implementations described below provide efficient anomaly detection by targeted (e.g., non-random) sampling of packets (e.g., in an ASIC PFE pipeline). This helps (multiple) collectors efficiently identify latency anomalies. The inventors have recognized that packet latency within a node can be affected by several factors. Key factors include, for example: (1) queuing delays caused by backpressure from various ASIC components / modules (e.g., switching structures, WAN (Wide Area Network) ports, etc.); (2) memory access latency (High Bandwidth Memory (HBM), Ternary Content Addressable Memory (TCAM), On-Chip Buffer (OCB), etc.); and / or (3) multiple recycling paths for packets if they undergo multiple ingress and / or egress pipeline processing. Queuing delays caused by backpressure from various ASIC components / modules are a major cause of latency. This is considered congestion within the port or ASIC. The example embodiments and implementations described below identify latency anomalies in the ASIC itself and use them when determining whether to sample, instead of randomly sampling packets, or sampling at a fixed preset frequency, or using some other factors (multiple) that do not take into account congestion within the node itself.
[0044] Some example embodiments consistent with this specification combine active queuing depth with congestion notification and / or queue priority to control the sampling of (multiple) packets, such that more important packets (e.g., packets with longer delays, longer normalized delays, and / or higher priority) are sampled with a higher probability than normal. This helps to obtain samples that are more affected by congestion. A collector receiving such samples from one or more nodes will have more candidate flows experiencing congestion and / or higher priority, and fewer candidate flows not experiencing congestion and / or not being high-priority flows. § 4.1 Example Network Topology for Deployment
[0045] Figure 4 The illustration shows an example network topology 400 that can be deployed consistent with the example embodiments described herein. Example network topology 400 includes a multicast video source server 410 that streams video data to client devices 460a and 460b. The video stream is transmitted via a client edge device (CE) 420a, a provider edge device (PE) 1 440a, one or more relay (core) routers 450a to 450d, PE 2 440b or PE 3 440c, and CE 2 420b or CE 3 420c. In the illustrated example, CE 3 420c multi-homed to the transport network via PE 2 440b and PE 3 440c. It is assumed that the links(multiple) with double-dotted lines and / or their associated routers are experiencing latency, and that the links(multiple) with three-line lines and / or their associated routers are not experiencing latency. Therefore, in this example, client device 1 460a will experience unwanted latency, while client device 2 460b will not. Thus, it is desirable to identify the root causes(s) of the latency. In the simplified diagram shown, only the relay (core) router 2 450b is shown providing out-of-band traffic monitoring information to collector 470, although other routers in the transport network 430 can (and are very likely) provide this information to collector 470. Collector 470 includes a Root Cause Analysis (RCA) module or engine 475. Figure 1 The same applies to the example. Figure 4 Some or all of the routers in the network will have agents (not shown). Furthermore, with... Figure 2 The same applies to the example. Figure 4 Some or all of the routers in the system have sampling modules / circuit systems in their respective ASIC packet pipelines and traffic monitoring modules in their respective software application blocks. § 4.2 (Multiple) Example Methods
[0046] Figure 5This is a flowchart of an example method 500 for performing packet processing and monitoring at a communication network (forwarding) node (e.g., a switch or router) in a manner consistent with this specification. As shown, the communication network node receives incoming packets. (Box 510) The packet forwarding engine (PFE) of the communication network node then queues the received incoming packets (e.g., in a Virtual Output Queue (VOQ)). (Box 520) (Review) Figure 3 Example method 500 then uses at least one of the following to generate a delay measurement or indication: (A) the queuing delay of the received inbound packet, and / or (B) the dwell time of the received inbound packet on at least a portion of the communication network node. (Box 530) In some examples, for some packets, this can be indicated (or inferred) by an explicit value or indicator or the absence of a delay indicator if there is no delay or no abnormal delay. Example method 500 then uses at least the generated delay measurement or indication to determine whether to sample the inbound packet. (Box 540) This can be performed by a sampling module / circuit system configured or adapted to perform such determination (recall, for example, Figure 2 Example method 500 then generates a message (e.g., not within a packet and neither appended to nor preceding a packet) for each of at least one sampled ingress packet, the message including at least one of the following: (i) absolute delay, (ii) expected delay associated with the communication network node, and / or (iii) cause of the anomaly (box 550), and sends the generated message toward the traffic sample collector for (e.g., root cause) analysis (box 560). These last two steps can be performed by a collector communication module specifically configured to execute the example method. (To recap, for example, Figure 2 (224.) Then leave example method 500. (Node 570)
[0047] Reference Figure 5 In box 530, in at least some example implementations, the dwell time of a received ingress packet at at least a portion of the communication network node is normalized as a function of the expected delay associated with the communication network node. In at least some of these example implementations, the expected delay associated with the communication network node is a function of at least one of the following: (A) the packet processing capability of the application-specific integrated circuit (ASIC) to which the ingress PFE belongs, (B) head-of-queue congestion at the queue, (C) the presence or absence of pass-through switching technology, (D) packet retrieval processing, (E) the type of on-chip memory used by the ASIC, (F) the type of off-chip memory used by the ASIC, (G) the port rate at which packets arrive at the port on the communication network node and / or (H) the port rate at which packets on the communication network node are forwarded to the port of the next hop.
[0048] Reference Figure 5 In box 550, in at least some example implementations, the generated message conforms to RFC 7011.
[0049] Referring again to box 550, the generated message can (and in most cases is) be separated from the sampled packets. In this way, the data packets themselves do not need to be modified (e.g., no header information needs to be added). In this case, referring to box 560, the action of sending the generated message toward the collector is separate from the action of forwarding the sampled packets to the next hop by the communication network nodes. That is, traffic monitoring information (e.g., sample information) is sent as a control message to the collector "out-of-band" or outside the data plane.
[0050] Still referring to box 550, in some example implementations, the cause of the exception is at least one of the following: (A) queuing delay at the VOQ, (B) memory access delay on the entry PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline through the entry PFE of the communication network node. To recap, for example, Figure 3 For example, in at least some example implementations, the entry PFE is part of an application-specific integrated circuit (ASIC) configured to monitor and provide at least one of the following: (A) queuing delay, (B) memory access delay on the entry PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline through the entry PFE of the communication network node.
[0051] Reference Figure 5 In box 540, at least in some example implementations, the action of whether to sample incoming packets is determined by the communication network node at least using a generated latency measurement or indication, or based on packet priority (e.g., based on service level, queue priority, etc.). In this example implementation, the probability of sampling higher-priority packets is higher than the probability of sampling lower-priority packets. (See reference...) Figure 3 Note that memory system 326 may include more than one VOQ for more than one output port and more than one priority (e.g., for more than one service class (CoS)).
[0052] In at least some example implementations of example method 500, the node's PFE can be used to perform each of the following operations: queuing incoming packets received in the Virtual Output Queue (VOQ); generating a delay measurement or indication; and determining whether to sample the incoming packets. In this example implementation, the PFE can be an application-specific integrated circuit (ASIC) or a part thereof. (To recap, for example, Figure 2 230.
[0053] Figure 6 This is a flowchart of an example method 600 for performing packet analysis at a collector in a manner consistent with this specification. As shown, the collector device receives multiple messages (e.g., IP flows conforming to RFC 7011) from multiple nodes in a communication network. (Box 610) Note that each of the multiple messages includes at least one of the following: (A) the absolute delay experienced by the flow at the node, (B) the expected delay of the node, and / or (C) the cause of the anomaly. (Review) Figure 5 (Box 550.) Further note that each of the multiple messages is separate from the data packets of the stream. Recall that the messages are generated / sent in the control plane, not the data plane. That is, these messages are out-of-band messages because they will likely propagate along different paths than the monitored data packets or streams. Next, example method 600 uses the received multiple messages to perform root cause analysis on the latency of one or more streams in the stream. (Box 620) For example, since the collector receives (or can determine) the normalized latency (the ratio of absolute latency to expected latency), this provides a degree of latency for comparing latency across two or more nodes. Otherwise, it would be difficult to compare congestion across different nodes based solely on absolute latency, if entirely possible. Furthermore, the collector can use information from the anomaly cause field to help determine the root cause of the stream latency. Then, we leave example method 600. (Node 630)
[0054] Referring to box 610, because example method 500 generates multiple samples (e.g., IP flow) included in messages in a non-random manner from each of multiple nodes, the collector is not overwhelmed by IP flow messages with unimportant sample information. Instead, the sample information sent to the collector for latency / delay analysis purposes has a higher relevance than randomly collected samples. § 4.3 Example Data Structures and / or Message Formats
[0055] Reference Figure 5 In box 550, at a given node, one or more samples can be collected and sent to the collector. In some example implementations consistent with this application, the format of these messages may be the same as or consistent with the message format described in RFC 7011 cited above. § 4.4 Example Apparatus
[0056] For example, a data communication network node can be a forwarding device, such as a router. Figure 7The illustration shows two data forwarding devices 710 and 720 coupled via a communication link 730. Link 730 can be a physical link or a “wireless” link. Data forwarding systems 710 and 720 can be, for example, routers. If data forwarding systems 710 and 720 are example routers, each data forwarding system may include control components (e.g., routing engines) 714 and 724 and forwarding components 712 and 722. Each data forwarding system 710 and 720 includes one or more interfaces 716 and 726 that terminate one or more communication links 730.
[0057] Reference Figure 2 In one example implementation, a sampling module / circuit system 232 suitable for performing sampling according to any of the described example embodiments may be provided in one of the forwarding components 712, 722. A sampling configuration module 222 and a collector communication module 224 suitable for and configured to perform sampling and message generation and transmission according to any of the described example embodiments may be provided in one of the control components 714, 724.
[0058] As just discussed above, and referring to... Figure 8 Some example routers 800 include a control component (e.g., a routing engine) 810 and a packet forwarding component (e.g., a packet forwarding engine) 890.
[0059] The control unit 810 may include an operating system (OS) kernel 820, multiple routing protocol processes 830, multiple label-based forwarding protocol processes 840, multiple interface processes 850, multiple user interface (e.g., command line interface) processes 860, and multiple chassis processes 870, and may store multiple routing tables 839, label forwarding information 845, and multiple forwarding (e.g., routing-based and / or label-based) tables 880. As shown, (multiple) routing protocol processes 830 may support routing protocols such as Routing Information Protocol (“RIP”) 831, Intermediate System to Intermediate System Protocol (“IS-IS”) 832, Open Shortest Path First Protocol (“OSPF”) 833, Enhanced Interior Gateway Routing Protocol (“EIGRP”) 834, and Border Gateway Protocol (“BGP”) 835, and (multiple) label-based forwarding protocol processes 840 may support protocols such as BGP 835, Label Distribution Protocol (“LDP”) 836, Resource Reservation Protocol (“RSVP”) 837, EVPN 838, and L2VPN 839. One or more components (not shown) may allow user 865 to interact with (multiple) user interface processes 860. Similarly, one or more components (not shown) may allow an external device to interact with one or more of the following processes via SNMP 885: router protocol process 830, label-based forwarding protocol process 840, interface process 850, and chassis process 870, and such processes may send information to the external device via SNMP 885.
[0060] The packet forwarding component 890 may include a microcore 892 on a hardware component (e.g., ASIC, switching structure, optics, etc.), (multiple) interface processes 893, an ASIC driver 894, (multiple) chassis processes 895, and (multiple) forwarding (e.g., routing-based and / or label-based) tables 896.
[0061] exist Figure 8In the example router 800, the control unit 810 handles tasks such as executing routing protocols, executing label-based forwarding protocols, and controlling packet processing. This frees up the packet forwarding unit 890 to quickly forward received packets. That is, received control packets (e.g., routing protocol packets and / or label-based forwarding protocol packets) are not fully processed by the packet forwarding unit 890 itself, but are passed to the control unit 810, thereby reducing the workload that the packet forwarding unit 890 must do and freeing it up to efficiently process packets to be forwarded. Therefore, the control unit 810 is primarily responsible for running routing protocols and / or label-based forwarding protocols, maintaining routing tables and / or label forwarding information, sending forwarding table updates to the packet forwarding unit 890, and performing system management. The example control unit 810 can process routing protocol packets, provide a management interface, provide configuration management, perform billing, and provide alerts. Processes 830, 840, 850, 860, and 870 can be modular and can interact with the OS kernel 820. That is, almost all processes communicate directly with the OS kernel 820. Modular software, which cleanly separates processes from each other, isolates problems in a given process so that such problems do not affect other processes that may be running. Additionally, using modular software facilitates easier scaling.
[0062] Still referencing Figure 8 The example OS kernel 820 may include an application programming interface (“API”) system for external program calls and scripting capabilities. The control unit 810 may be based on an Intel PCI platform running the OS from flash memory, with alternating copies stored on the router's hard drive. The OS kernel 820 is layered on the Intel PCI platform and establishes communication between the processes of the Intel PCI platform and the control unit 810. The OS kernel 820 also ensures that the forwarding table 896 used by the packet forwarding unit 890 is synchronized with the forwarding table 880 in the control unit 810. Therefore, in addition to providing the underlying infrastructure to the software processes of the control unit 810, the OS kernel 820 also provides a link between the control unit 810 and the packet forwarding unit 890.
[0063] refer to Figure 8Multiple routing protocol processes 830 provide routing and route control functions within the platform. In this example, RIP 831, ISIS 832, OSPF 833, and EIGRP 834 (and BGP 835) protocols are provided. Of course, other routing protocols may be provided additionally or alternatively. Similarly, multiple label-based forwarding protocol processes 840 provide label forwarding and label control functions. In this example, LDP 836, RSVP 837, EVPN 838, and L2VPN 839 (and BGP 835) protocols are provided. Of course, other label-based forwarding protocols (e.g., MPLS, SR, etc.) may be provided additionally or alternatively. In example router 800, multiple routing tables 839 are generated by the multiple routing protocol processes 830, while label forwarding information 845 is generated by the multiple label-based forwarding protocol processes 840.
[0064] Still referencing Figure 8 The (multiple) interface processes 850 execute the configuration of physical interfaces and encapsulation.
[0065] Example control unit 810 can provide several ways to manage the router. For example, example control unit 810 can provide multiple user interface processes 860 that allow system operator 865 to interact with the system through configuration, modification, and monitoring. SNMP 885 allows SNMP-enabled systems to communicate with the router platform. This also allows the platform to provide necessary SNMP information to external agents. For example, SNMP 885 can allow the system to be managed from a network management station running software such as Hewlett-Packard's Network Node Manager (“HP-NNM”) through a framework such as Hewlett-Packard's OpenView. Packet billing (often referred to as traffic statistics, and which includes managing / configuring packet sampling settings, consumption, analysis, and / or reporting of sampling results, etc., according to any example methods described in this application) can be performed by control unit 810, thereby avoiding slowing down traffic forwarding by packet forwarding unit 890.
[0066] Although not shown, the example router 800 may provide out-of-band management, an RS-232 DB9 port for serial console and remote management access, and three-tier storage using a removable PC card. Furthermore, although not shown, a process interface located at the front of the chassis provides an external view into the router's internal operations. It can be used as a troubleshooting tool, a monitoring tool, or both. The process interface may include LED indicators, alarm indicators, control component ports, and / or displays. Finally, the process interface may provide interaction with the command-line interface (“CLI”) 860 via a console port, auxiliary port, and / or management Ethernet port.
[0067] The packet forwarding unit 890 is responsible for outputting received packets correctly as quickly as possible. If there is no entry for a given destination or a given label in the forwarding table, and the packet forwarding unit 890 cannot perform forwarding on its own, it may send packets destined for that unknown destination to the control unit 810 for processing. The example packet forwarding unit 890 is designed to perform Layer 2 and Layer 3 switching, route lookup, and fast packet forwarding.
[0068] like Figure 8 As shown, the example packet forwarding unit 890 has an embedded microcore 892 on hardware unit 891, multiple interface processes 893, ASIC driver 894, and multiple chassis processes 895, and stores multiple forwarding (e.g., route-based and / or tag-based) tables 896. The microcore 892 interacts with the multiple interface processes 893 and multiple chassis processes 895 to monitor and control these functions. The multiple interface processes 892 have direct communication with the OS kernel 820 of control unit 810. This communication includes forwarding abnormal packets and control packets to control unit 810, receiving packets to be forwarded, receiving forwarding table updates, providing control unit 810 with information about the health of packet forwarding unit 890, and allowing configuration of interfaces from multiple user interface (e.g., CLI) processes 860 of control unit 810. The stored multiple forwarding tables 896 are static until a new forwarding table is received from control unit 810. Multiple interface processes 893 use multiple forwarding tables 896 to look up next-hop information. Multiple interface processes 893 also have direct communication capabilities with the distributed ASIC. Finally, multiple chassis processes 895 can communicate directly with the microcore 892 and with the ASIC driver 894.
[0069] Reference Figure 2 In one example implementation, the sampling module / circuit system 232, suitable for performing sampling according to any of the example embodiments described, can be... Figure 8 The packet forwarding unit 890 is provided in hardware components 891, microcore 892, ASIC driver 894, and / or interface process 893. It can be... Figure 8 The control unit 810 provides a sampling configuration module 222 and a collector communication module 224 in its flow monitoring process or daemon (not shown), which are adapted to and configured to perform sampling and message generation and sending according to any example embodiment of the described example embodiments. Further, the module is adapted to perform sampling according to any example embodiment of the described example embodiments. Figure 3 Example Packet Forwarding Engine (PFE) 300 can be used Figure 8 The packet forwarding component 890 is provided.
[0070] Although it is possible Figure 7 or Figure 8 The example router described herein implements the example embodiments consistent with this specification, but the embodiment consistent with this specification can be implemented on communication network nodes (e.g., routers, switches, etc.) with different architectures. More generally, the embodiment consistent with this specification can be implemented on... Figure 9 The example system 900 illustrated is implemented on this system.
[0071] Figure 9 This is a block diagram of an exemplary machine 900 that can execute one or more processes described in the process and / or store information used and / or generated by such processes. The exemplary machine 900 includes one or more processors 910, one or more input / output interface units 930, one or more storage devices 920, and one or more system buses and / or networks 940 for facilitating information communication among the coupled elements. One or more input devices 932 and one or more output devices 934 may be coupled to one or more input / output interfaces 930. The one or more processors 910 can execute machine-executable instructions (e.g., C or C++ running from a Linux operating system widely available from many vendors) to implement one or more aspects of this specification. At least a portion of the machine-executable instructions may be (temporarily or more permanently) stored on one or more storage devices 920 and / or may be received from an external source via one or more input interface units 930. The machine-executable instructions may be stored as various software modules, each performing one or more operations. Functional software modules are examples of components of this specification.
[0072] In some embodiments consistent with this specification, processor 910 may be one or more microprocessors and / or ASICs. Bus 940 may include a system bus. Storage device 920 may include system memory, such as read-only memory (ROM) and / or random access memory (RAM). Storage device 920 may also include a hard disk drive for reading from and writing to a hard disk, a disk drive for reading from or writing to a (e.g., removable) disk, an optical disk drive for reading from or writing to a removable (magneto)optic disk (such as a compact disc or other (magneto)optic media), or a solid-state non-volatile storage device.
[0073] Some exemplary embodiments consistent with this specification may also be provided as machine-readable media for storing machine-executable instructions. Machine-readable media may be non-transitory and may include, but is not limited to, flash memory, optical disc, CD-ROM, DVD-ROM, RAM, EPROM, EEPROM, magnetic cards or optical cards, or any other type of machine-readable media suitable for storing electronic instructions. For example, exemplary embodiments consistent with this specification may be downloaded as a computer program that can be transmitted from a remote computer (e.g., a server) to a requesting computer (e.g., a client) via a communication link (e.g., a modem or network connection) and stored on a non-transitory storage medium. Machine-readable media may also be referred to as processor-readable media.
[0074] The exemplary embodiments (or components or modules thereof) consistent with this specification may be implemented in hardware such as one or more field-programmable gate arrays (“FPGAs”), one or more integrated circuits (such as ASICs), one or more network processors, etc. Alternatively or additionally, the embodiments (or components or modules thereof) consistent with this specification may be implemented as stored program instructions executed by a processor. Such hardware and / or software may be provided in addressing data (e.g., packets, cells, etc.) forwarding devices (e.g., switches, routers, etc.), laptop computers, desktop computers, tablet computers, mobile phones, or any device with computing and networking capabilities. § 4.5 Refinement, Alternatives and / or Expansion § 4.5.1 Example Delay / Latency Measurement or Indicator
[0075] Reference Figure 3 Box 530, most (if not all) ASICs in today's routing industry can measure packet dwell time by timestamping the arrival and departure of packets on (multiple) ports. Such ASICs can also monitor queued traffic using various threshold settings. § 4.5.1.1 Example of determining “normalized” delay / delay
[0076] Packet dwell time in a node's ASIC pipeline refers to the time a packet has spent in the ASIC. Packet dwell time can depend on factors such as: - Packet processing capabilities of ASICs. For example, the packet processing latency of an older generation ASIC with 80 Gbps will be significantly different from that of a newer generation ASIC with 1.2 Tbps; - Queuing and internal congestion delays at the front of the queue; - Cut-through switching technology; - Grouped recycling processing; - (Multiple) memory types (e.g., HBM, TCAM, SRAM, OCB, etc.); and / or - Port speed. Due to these factors(s), the absolute dwell time of a packet can vary depending on the ASIC design. Therefore, absolute delay or absolute dwell time may not accurately represent the relative congestion encountered by packets on the device.
[0077] At least some example embodiments employ a normalization technique that allows for improved grading of congestion levels within a node. In one example, the minimum expected latency for a packet is defined as 0 based on device characteristics. For instance, a 1 μs dwell time / latency might be considered normal (i.e., no congestion) on an 80 Gbps packet processing ASIC, but the same 1 μs dwell time / latency might be considered abnormal on a 1.2 Tbps packet processing ASIC. In some example implementations, nodes share their expected dwell time for packets along with the measured value. Alternatively, nodes can identify their brand and model, and the collector can look up their expected dwell time. The expected dwell time can be pre-calculated based on manufacturer RFC2544 testing and can be used in the node software.
[0078] The ratio of measured (i.e., actual) stay time to expected stay time provides information about the severity of the delay (abnormality). § 4.5.2 (Multiple) Example Sampling Determination
[0079] Referring to box 540, the action of determining whether to sample the ingress group, using at least the generated delay measurement or indication, can be accomplished in several different alternatives. The sampling decision can use one or more of the following: - Conditionally, based on a delay measurement value exceeding a threshold; - Conditionally, based on normalized delay measurements exceeding a threshold; - Conditionally, based on whether a delay indicator is set on the node (e.g., the node's PFE ASIC); - Conditionally, based on grouping priority being higher than a threshold; - Conditional, based on the group having a specific service category (CoS); - Probabilistic, based on delay measurements (where the sampling probability increases with the value of the delay measurement); - Probabilistic, based on normalized delay measurements (where the sampling probability increases with the value of the normalized delay measurement); - Probabilistic, based on whether the node has a delay indicator set (where the sampling probability increases if a delay indicator is set); - Probabilistic, based on group priority (where the sampling probability increases with group priority); and / or - Probabilistic, group-based CoS (where the sampling probability increases with the CoS of the group). Thresholds are typically preset or pre-configured, but can also be dynamically determined using a pre-configured algorithm. Additionally, other factors can be considered. Furthermore, two or more sampling determinations can be performed, where one or more of these determinations consider any combination of the factors(s) listed above, and where one or more of these determinations consider (i.e., factors unrelated to grouping delay or priority). Such sampling determinations can be performed independently (e.g., in this case, the determination to sample via any one of the determinations will cause sampling to occur) or dependently (e.g., in this case, sampling only occurs when all sampling determinations have determined sampling). That is, in the former case, the result of a single sampling determination is logically OR, while in the latter case, the result of a single sampling determination is logically AND. A non-exhaustive discussion of an example sampling decision process is now provided.
[0080] As an example, whether to sample a packet can be simply based on the presence or absence of a received delay indication (e.g., a delay bit set by the ASIC). That is, if a delay is indicated, then (e.g., only) the packet is sampled. Packet priority can also be considered. For example, if a delay condition is met and its priority (e.g., CoS, queue priority, etc.) is above a certain threshold, then (e.g., only) the packet is sampled.
[0081] As another example, the probability of sampling a group is a function of the delayed measurement or the normalized delayed measurement. For instance, the probability of sampling a group may increase as the delayed measurement or the normalized delayed measurement increases. As yet another example, whether or not a group is sampled may be subject to the delayed measurement or the normalized delayed measurement exceeding a threshold value. Group priority, as discussed above, can also be considered.
[0082] In some example embodiments, a communication network node may perform packet flow sampling based on at least one of the following: (A) packet priority (e.g., based on service level, queue priority, etc.) and / or (B) the delay length at the queue at the ingress packet forwarding engine (PFE) of the communication network node.
[0083] In some example embodiments, packet streams are sampled in a non-random manner such that packet streams experiencing local anomalous delays are sampled more frequently than another packet stream that does not experience local anomalous delays, and / or higher-priority packet streams are sampled more frequently than lower-priority packet streams. § 4.5.3 Example Sample Reporting Technology
[0084] Referring to return box 550, the example embodiments consistent with this specification may provide new information (e.g., as a new field in a message conforming to RFC 7011) that will help the collector easily and accurately identify anomalous flows. This new information may include one or more of, for example, the following: - The absolute dwell time of the group; - Minimum expected stay time for each group; - Normalized dwell time for grouping; and / or - Abnormal cause (e.g., queuing delay). Using the cause of the congestion, the collector can determine that the flow experienced congestion in a node due to (A) queuing delay, (B) (multiple) recycling paths, and / or (C) any other reason. Using dwell time information, the expected minimum dwell time for a group, and / or the normalized dwell time for a group, or one or more of these, the collector can understand the severity of the congestion. §4.5.4 Example Implementations of Moving at Least Some Functionality from Software Application Blocks (e.g., Control Plane) to ASIC Packet Pipelines (e.g., Forwarding Plane)
[0085] Reference Figure 2 In some alternative implementations, at least some of the functionalities of the software application block 220 can be performed by the ASIC packet pipeline 230. As an example, some functionalities of the collector communication module 224 can also be offloaded to the ASIC. For example, a collector configuration module (not shown) can be used to program the IP address and UDP port of the collector (server) 290 in the ASIC 230. In this example implementation, the ASIC 230 can be provided with logic to (1) buffer samples, (2) generate IPFIX messages, and / or (3) send IPFIX messages directly to the remote collector (server) 290 without sending them to the software application block 220. This offloading of IPFIX message generation and / or transmission to hardware is sometimes referred to as “inline IPFIX”. § 4.6 Conclusion
[0086] The example embodiments consistent with this specification can provide one or more of the following advantages. First, since latency-based anomaly detection is local to the node (“box”), the collector’s workload is reduced. Conversely, most of the workload is distributed across nodes. Second, since delayed flows will likely be sampled more frequently, the information reported to the collector is more relevant and therefore more efficient. For example, sampling frequency can be improved by taking into account the flow experiencing congestion and / or queue priorities. Third, by providing the collector with “normalized” latency, the collector can classify and take corrective action on the latency across (multiple) end-to-end flows. Fourth, these example methods are easier to implement using existing router ASICs. Fifth, anomalies can be detected in real time without affecting the data flow due to the addition of (multiple) headers or probes. The example embodiments allow flow monitoring to be selectively applied to queues (e.g., based on latency and / or priority). Congestion bits, which are becoming increasingly available in modern ASICs, can be used for anomaly detection. Reporting has been improved (e.g., through enhancements to the IPFIX standard template) to share congestion indicators, normalized latency, and / or dwell time with the collector.
Claims
1. A method for use on a communication network node, the method comprising: a) The ingress packet is received by the communication network node; b) The incoming packets received in the Virtual Output Queue (VOQ) are queued by the Packet Forwarding Engine (PFE) of the communication network node; c) The communication network node generates a delay measurement or indication based on at least one of the following: (A) the queuing delay of the received inbound packet, and / or (B) the dwell time of the received inbound packet on at least a portion of the communication network node; d) The communication network node determines whether to sample the ingress packet using at least the generated delay measurement or indication; e) A message generated by the communication network node, for example, not within the packet and neither appended to nor preceding the packet, for each sampled ingress packet in at least one sampled ingress packet, the message including at least one of the following: (i) absolute delay, (ii) expected delay associated with the communication network node, and / or (iii) cause of the anomaly; and f) The generated message is sent by the communication network node toward the traffic sample collector for analysis, such as root cause analysis.
2. The method of claim 1, wherein the dwell time of the received ingress packet on at least a portion of the communication network node is normalized as a function of the expected delay associated with the communication network node.
3. The method of claim 2, wherein the expected delay associated with the communication network node is a function of at least one of the following: (A) the packet processing capability of the application-specific integrated circuit (ASIC) to which the ingress PFE belongs, (B) head-of-queue congestion at the queue, (C) the presence or absence of pass-through switching technology, (D) packet retrieval processing, (E) the type of on-chip memory used by the ASIC, (F) the type of off-chip memory used by the ASIC, (G) the port speed at which the packet arrives at the port on the communication network node and / or (H) the port speed at which the packet on the communication network node is forwarded to the port of the next hop.
4. The method of claim 1, wherein the generated message conforms to RFC 7011.
5. The method of claim 1, wherein the generated message is separate from the sampled packets, and The action of sending the generated message toward the collector is separate from the action of the communication network node forwarding the sampled packet to the next hop.
6. The method of claim 1, wherein the communication network node determines whether to sample the ingress packet based on packet priority or service level CoS, at least using the generated delay measurement or indication.
7. The method of claim 1, wherein the cause of the anomaly is at least one of the following: (A) queuing delay at the VOQ, (B) memory access delay on the ingress PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline of the ingress PFE of the communication network node.
8. The method of claim 7, wherein the entry PFE is part of an application-specific integrated circuit (ASIC) configured to monitor and provide at least one of the following: (A) the queuing delay, (B) the memory access delay on the entry PFE of the communication network node, and / or (C) the one or more recycling paths of the pipeline through the entry PFE of the communication network node.
9. The method of claim 1, wherein the following actions are performed by the PFE: - Queue the incoming packets received in the virtual output queue (VOQ). - Generate delay measurements or indications, and - Determine whether to sample the ingress group.
10. The method of claim 9, wherein the PFE is an application-specific integrated circuit (ASIC) or a portion thereof.
11. A computer-implemented method, comprising: a) The collector device receives multiple flow congestion messages from multiple nodes in the communication network. Each of the plurality of flow congestion messages includes at least one of the following: (A) the absolute delay experienced by the flow at the node, (B) the expected delay for the node, and / or (C) the cause of the anomaly. Each of the multiple flow congestion messages is separate from the data packets of the flow; as well as b) Use the received multiple IP stream messages to perform a root cause analysis of the latency of one or more streams in the stream.
12. The computer-implemented method of claim 11, wherein the samples included in the plurality of flow congestion messages are generated in a non-random manner by each of the plurality of nodes.
13. A communication network node, comprising: a) Interface, suitable for receiving incoming packets; b) Packet Forwarding Engine (PFE), suitable for 1) Queuing the incoming packets received in the virtual output queue (VOQ), and 2) Generate a delay measurement or indication based on at least one of the following: (A) queuing delay of an incoming packet received by the communication network node, and / or (B) the dwell time of the incoming packet received by the communication network node on at least a portion of the communication network node; as well as c) An application-specific integrated circuit (ASIC) suitable for determining whether to sample the ingress packet, using at least the generated delay measurement or indication.
14. The communication network node according to claim 13, further comprising: d) At least one processor is configured to 1) Generate a message for each sampled ingress packet in at least one sampled ingress packet, the message including at least one of the following: (i) absolute delay, (ii) expected delay associated with the communication network node, and / or (iii) cause of the anomaly; and 2) Send the generated message toward the traffic sample collector for root cause analysis.
15. The communication network node of claim 14, wherein the cause of the anomaly is at least one of the following: (A) queuing delay at the VOQ, (B) memory access delay on the ingress PFE of the communication network node, and / or (C) one or more recycling paths in the pipeline of the ingress PFE of the communication network node.
16. The communication network node of claim 13, wherein the dwell time of the received ingress packet on at least a portion of the communication network node is normalized as a function of an expected delay associated with the communication network node.
17. The communication network node of claim 16, wherein the expected delay associated with the communication network node is a function of at least one of the following: (A) the packet processing capability of the application-specific integrated circuit (ASIC) to which the ingress PFE belongs, (B) head-of-queue congestion at the queue, (C) the presence or absence of pass-through switching technology, (D) packet recycling processing, (E) the type of on-chip memory used by the ASIC, (F) the type of off-chip memory used by the ASIC, (G) the port speed at which the packet arrives at the port on the communication network node and / or (H) the port speed at which the packet on the communication network node is forwarded to the port of the next hop.
18. The communication network node of claim 13, wherein the generated message conforms to RFC 7011.
19. The communication network node of claim 13, wherein the generated message and the sampled packet are separate, and The action of sending the generated message toward the collector is separate from the action of the communication network node forwarding the sampled packet to the next hop.
20. The communication network node of claim 13, wherein when at least the generated delay measurement or indication is used to determine whether to sample the ingress packet, the ASIC is adapted to also consider packet priority, for example, based on service level, queue priority, etc.