Network error reporting and recovery with inline packet processing pipelines

By integrating hardwired logic circuits in the switch's packet processing pipeline, link statistics are processed in real time and alarm messages are generated, solving the problem of throughput degradation caused by packet loss or damage in the network, and achieving fast recovery and efficient network management.

CN120712759APending Publication Date: 2025-09-26INTEL CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380095219.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-04-03
Filing Date
2023-11-10
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

As network speeds increase, the frequency of packet loss or damage also increases, affecting overall throughput. Traditional per-stream retransmission mechanisms have long recovery times, and link statistics collection and recovery intervention functions are slow, making them unable to promptly address packet loss or damage issues.

Method used

By integrating dedicated hard-wired logic circuits in the packet processing pipeline between the MAC circuit module and the switch core, link statistics are processed in real time, and packets containing link telemetry information and alarm messages are generated and sent directly to the endpoints of the flow or the network management system to achieve a fast recovery process.

Benefits of technology

It reduces the recovery time of packet loss or damage, improves the response speed and throughput of the network, reduces the performance degradation caused by packet loss, and enhances the real-time and efficiency of network management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120712759A_ABST
    Figure CN120712759A_ABST
Patent Text Reader

Abstract

An apparatus is described. The apparatus includes an electronic circuit module to support a plurality of streams within the network. The electronic circuitry module is used to determine respective telemetry information for the plurality of streams and inject an alarm message into a particular one of the plurality of streams when an alarm condition is reached for the particular one of the streams. The alert message includes a multi-bit error code describing an alert condition. The multi-bit error code is one of a plurality of possible multi-bit error codes.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Declaration This application claims priority under 35 U.S.C. § 365(c) to U.S. Application No. 18 / 130,383, filed April 3, 2023, which is hereby incorporated by reference in its entirety. Background Art

[0002] As computing environments continue to rely on high-speed, high-bandwidth networks to interconnect their various computing components, system administrators are becoming increasingly concerned about the tendency of networks to lose more information as their performance increases. BRIEF DESCRIPTION OF THE DRAWINGS

[0003] Figure 1a illustrates a high performance computing environment; Figure 1b shows a networking switch; Figure 2 An improved networking switch is shown; Figure 3a 、 Figure 3b and Figure 3c Involving the use of Figure 2 Improved error reporting and recovery mechanisms implemented by networking switches; Figure 4 Involving the use of Figure 2 Another error reporting and recovery mechanism implemented by the improved networking switch; Figure 5 Describe another high performance computing system; Figure 6a and Figure 6b Describe the IPU; Figure 7 Describe the packet processing pipeline. DETAILED DESCRIPTION

[0004] Figure 1a A high performance computing environment 100, such as a data center, is shown. Figure 1a As viewed in FIG, a high performance computing environment 100 includes multiple units of high performance computing equipment (e.g., a rack-mounted CPU unit 101, a rack-mounted memory unit 102, and a rack-mounted storage device unit 103) that are communicatively coupled to a network 104. The high performance computing equipment 101, 102, 103 transmit data packets and / or commands between each other over the network 104.

[0005] As the end-to-end propagation delay of packets through network 104 decreases, the overall performance of computing environment 100 is improved (receiving devices receive their incoming packets faster and can therefore act on the contents of the packets faster). However, the problem is that as the speed of network 104 increases, the propensity of network 104 to corrupt or lose packets also increases.

[0006] Traditionally, lost packets have been handled through various per-flow retransmission mechanisms. Here, a flow is a unique logical "connection" between two endpoints (an endpoint can be a unit of a high-performance device, or a component within such a device, such as a CPU core within a multi-core CPU processor) over network 104. Typically, each flow is at least partially defined by a unique combination of source and destination addresses (other information, such as the applicable protocol, may also define a flow). At any given moment, a network typically supports a large number of flows, reflecting the number of distinct pairs of device endpoints in communication sessions with each other.

[0007] According to traditional streaming protocols, the sending endpoint does not remove a packet from its local memory until the receiving endpoint has acknowledged its receipt of the packet. If the sending endpoint does not receive an acknowledgement for a packet (or sequence of packets), the sending endpoint resends the packet(s) to the destination endpoint.

[0008] The problem is that as the frequency of lost or damaged packets along a particular flow increases, the overall throughput of that flow suffers. Here, the recovery time for lost / damaged packets is quite long because the sending endpoint must wait for a considerable predetermined amount of time (timeout) before it resends the lost / damaged packet if no acknowledgment is received.

[0009] Network nodes within the network may also monitor packet loss / corruption statistics and attempt to intervene (e.g., reroute the connection). Figure 1b As observed in

[15] , statistics collection and recovery intervention functions are typically implemented as centralized, slower software routines.

[0010] Figure 1b An on-chip switch architecture 120 is shown, which includes an ingress packet processing pipeline 123 between an ingress media access control (MAC) layer circuit module 122 and a switch core 124. In the ingress direction, a plurality of ingress links 121 feed into the ingress MAC circuit module 122, which in turn feeds into the ingress packet processing pipeline 123. The ingress MAC circuit 122 controls / supervises the inbound links 121 and passes received packets to the ingress pipeline 123.

[0011] The ingress packet processing pipeline 123 processes ingress packets and forwards them to the switch core 124. Based on, for example, the packet's corresponding destination address, the packet is routed to the appropriate egress path, which includes the egress packet processing pipeline 225, MAC layer circuitry 226, and corresponding egress links 227. The egress packet processing pipeline 125 constructs the IP header fields for the outbound packet. The egress MAC circuitry module 126 appends link layer header information to the packet and physically sends each of these packets over one of the egress links 127.

[0012] While link statistics are tracked separately for ingress link 121 and egress link 127 at the media access control (MAC) layer 122, the statistics are collected by polling 131 various registers within the MAC layer circuitry 122 (for ease of illustration, only the ingress-side polling 131 is depicted). Here, a general-purpose processor core 132 executes polling software that (e.g., in a round-robin fashion) accesses each statistics register for each link individually and then stores the collected data in memory 133. Reading the statistics registers is inherently a slow, serial data collection process.

[0013] After being stored in memory 133 , the data is then analyzed by software running on processing core 132 , which causes processing core 132 to execute hundreds or thousands of instructions (or more) to analyze the data.

[0014] If a problem is observed in one of the links (eg, too many errors along a particular link), the processing core 132 sends a notification to the central management system 105 of the network (see back for details). Figure 1a ) sends an alert 134. In response, the central management system 105 triggers the start of some recovery algorithm (eg, rerouting the affected flows to avoid the problematic link).

[0015] Here, after the MAC layer statistics reveal a problem, serial data collection and data analysis in the software consume tens or hundreds of milliseconds before recovery from the problem is initiated. The tens or hundreds of milliseconds spent before an alarm ("ALARM") signal is generated may result in many packets being dropped between the time the MAC circuit module 122 first generates an error message and the time any corrective action is implemented.

[0016] refer to Figure 2One solution is to instead forward the link statistics information 233 from the MAC circuit module 222 to the ingress and / or egress packet processing pipelines 223, 226, and design the packet processing pipelines 223, 226 to construct packets containing link telemetry information and / or containing alarm messages derived from the link telemetry information. The link telemetry information can include link statistics ("link statistics") information, information derived from the link statistics information, or any combination of such information. Notably, part of the packet construction process includes inserting destination address information within the packet header that specifies the endpoints (e.g., destination endpoints) of one or more of the flows currently flowing through the system 220 and / or the network management system.

[0017] By constructing such packets on the fly and having them received by one or more flow endpoints and / or a network management system shortly thereafter, the recovery process may be initiated shortly after the MAC circuit module 222 generates link statistics necessitating recovery.

[0018] Here, by implementing the packet processing pipelines 223 and 236, for example, by using dedicated hard-wired logic circuit modules integrated on the same semiconductor chip 220 as the MAC circuit module 222 that collects link statistics, these pipelines 223 and 236 can process the link statistics in hardware almost immediately after the MAC circuit module 222 first generates the information. Therefore, time-consuming serial polling of link statistics from the MAC circuit module 222 and processing of the link statistics in software can be avoided.

[0019] Figure 3a 、 Figure 3b and Figure 3c Different methods for constructing and sending packets containing link telemetry information and / or alert messages as described above are described.

[0020] Figure 3a Concerning the first method, the ingress and / or egress packet processing pipelines 223, 226 immediately generate a packet containing an alert message upon receiving header information of a packet determined to be damaged from the ingress MAC circuitry module 222 downstream from the ingress MAC circuitry module 222. Here, for example, when a damaged ingress packet is detected by the ingress MAC circuitry 222 at 301, if the header information of the packet is valid (e.g., the damage is within the packet payload), the MAC circuitry 222 forwards the header and optionally a portion of the packet's payload to the ingress packet processing pipeline 223, along with, for example, a specific error code (among a plurality of possible error codes) corresponding to the exact error identified by the MAC unit.

[0021] Then, at 302, the packet processing pipeline 223 uses the forwarded header and error code to construct an alert packet. This alert packet can be sent to the source endpoint of the packet's flow, the destination endpoint of the packet's flow, or both, to notify the endpoint(s) of the error. The endpoint(s) can then initiate recovery procedures (e.g., retransmit the packet (source endpoint)), request retransmission of the packet (the destination endpoint sends a request to the source endpoint), and / or issue an alert to the network management system. At 304, the damaged original packet is cleared / discarded by the MAC circuitry 222 or the packet processing pipeline.

[0022] Alternatively, metadata of the damaged packet can be set (e.g., by MAC circuitry 222 or ingress pipeline 223) to indicate that the packet is damaged. The data structure representing the packet is then switched via switch 224 to the correct egress pipeline 225 (e.g., if the alarm message is to be sent to a destination endpoint, the egress pipeline 225 is associated with the destination endpoint). The egress pipeline 225 observes from the metadata (which is logically attached to the data structure) that the packet is damaged, generates an alarm message, and sends the alarm message to the destination endpoint of the flow of damaged packets, for example.

[0023] It is noteworthy that the alert message may include a multi-bit error code that specifies a particular problem (i.e., a determination that the payload of a packet is damaged). Here, the particular multi-bit error code is selected from among a plurality of possible multi-bit error codes (e.g., multiple bits are required to express multiple different problems).

[0024] If the header information of the packet is invalid, the damaged header is forwarded to the ingress packet processing pipeline 223, which stores it in local memory. Then, at 303, the pipeline 223 appends the damaged header as an additional payload to any / all subsequent packets that are not damaged and are processed by the pipeline 223. Each such packet having an additional payload with a damaged header may include an alert message with another multi-bit error code that specifies the specific problem, i.e., another packet that may belong to the same flow as the current packet (which is carrying the additional payload) is believed to have a damaged header.

[0025] Ideally, such packets arrive at their destination endpoint, which handles the multi-bit error code and the damaged header that has been included as additional payload. Each receiving endpoint determines from the contents of the damaged header whether the packet with the damaged header is likely from its particular flow. If any receiving endpoint makes such a determination, it can trigger recovery with the sending endpoint (requesting the packet to be resent) and / or send an alert to the network management system.

[0026] In addition to sending an alert message to the network management system from the destination endpoint of a flow whose packets are known to be corrupted or believed with high confidence to be corrupted, the ingress packet processing pipeline 223 that receives invalid header information can also construct an alert message that includes the corrupted header information and a multi-bit error code and send it directly to the network management system.

[0027] Furthermore, even if neither the MAC circuitry module 222 nor the packet processing pipeline 223 can determine whether the packet header is corrupted, the packet processing pipeline 223 can create an alert message with the packet's header information and send it to either or both of the packet's source and destination endpoints, allowing those endpoints to determine whether the packet header is corrupted. (The alert message can be a separate packet from the packet with the header whose corruption status is uncertain, or included within (e.g., appended to) the packet with such a header.) The alert message can include another multi-bit error code that specifies the problem, namely, that the header corruption status is unknown. The multi-bit error code (or data associated with the error code) can include information regarding the lack of knowledge regarding possible corruption to ensure that the packet can be safely discarded (e.g., in the case of corruption in the source / destination address itself). Upon receiving the alert message, the source / destination endpoint(s) can match the packet header fields against their active connections. If there is no match, the packet header is corrupted. If there is a match, it is likely that the packet header is not corrupted.

[0028] In the above embodiment, it is noted that the network addresses of the destination and / or source endpoints of the flow need not (but may) be explicitly identified in any packet carrying the alert message. Here, consistent with label switching or other flow processes that change the source and / or destination header information of the packet, the switching / routing function of the switch directs the packet to the correct egress port for the flow of the packet.

[0029] In various embodiments, the network management system is at least partially distributed across the network's constituent switching nodes, including the packet processing pipeline's own switching node, in which case the packet processing pipeline sends internal communications only to software executing locally on that switching node. Alternatively, the packet processing pipeline may incorporate a destination address of an external network node designated for the network management system into the header of the alert message packet.

[0030] Figure 3b A method is involved in which, for example, network telemetry information collected at each node hop of a flow 311 through a network 304 is sent to a destination endpoint 312 of the flow. Figure 3b As observed in , flow 311 flows from source endpoint 312 through switches A, B, and C to destination endpoint 312.

[0031] Here, the ingress MAC circuit module 222 and / or the ingress pipeline 223 of the ingress link of switch A that receives packets of flow 311 collects telemetry information ("A statistics") for the link. To name just a few possibilities, the link telemetry information may include any of the following: 1) the total error count since a global counter reset (e.g., resetting all link error counters in the network to 0); 2) the error count within a recent time window (where the time window is shortened and continuously repeated); 3) with a timestamp in #1) above; 4) with a timestamp in #2) above; 5) with a link ID in #1) or 3) above; 6) with a link ID in #2) or 4) above, etc.

[0032] When a packet of flow 311 is received from an ingress link at switch A, telemetry information for that link is collected by the MAC circuit module 222 and / or the ingress packet processing pipeline 223, and then processed by either or both of switch A's packet processing pipelines 223 and 225. Either or both of the packet processing pipelines 223 and 225 construct a header for the packet containing the link's telemetry information (alternatively, the link's telemetry information can be appended to the packet as additional payload). The packet is then transmitted from the first switch A to the second switch B.

[0033] Similarly, telemetry information for the ingress link of switch B that receives the packet is continuously collected by the MAC circuit module 222 and / or the ingress packet processing pipeline 223 of switch B and then processed by either or both of the packet processing pipelines 223, 225 of switch B. When a packet is received by switch B and then processed by either or both of the pipelines 223, 225 within switch B, the pipeline(s) construct header information for the packet that somehow accumulates or combines the telemetry information for the ingress link to switch A (which is carried by the packet from switch A to switch B) and the telemetry information for the ingress link to switch B. Figure 3b The accumulated error statistics are described as "A+B statistics".

[0034] In the basic approach, link telemetry information counts the total errors at each link, and accumulation adds the two counts from two links to produce a single total error count (scalar). In another approach, accumulation lists the corresponding error counts for the two links as two different numbers (vectors). In either of these approaches, the error count can be the total error count (e.g., since a global reset) or the error count within the most recent time window (which is reset to zero after each time window expires). For any of these approaches, along with the error statistics for the specific link whose telemetry is incorporated into the packet, a timestamp and / or ID of the link can also be included.

[0035] In any case, after the accumulated telemetry information (A+B statistics) has been integrated into the packet, the packet is transmitted from the second switch B to the third switch C along flow 311.

[0036] The process is then repeated for the third switch C, with the result that the telemetry information accumulated for the three respective ingress links into switches A, B, and C (“A+B+C statistics”) is incorporated into the packet before it is sent from the third switch C to the receiving endpoint 313.

[0037] Destination endpoint 313 can then process the telemetry information to determine whether there is a problem along packet flow 311 and, if so, flag an error. For example, if the telemetry information is presented as a scalar (errors across all three links are summed), endpoint 313 can use a predetermined threshold to determine whether there is a problem (e.g., if the scalar count exceeds the threshold, then there is a problem). As another example, if the telemetry information is presented as a vector (errors from all three links are provided separately), endpoint 313 can use a lower threshold predetermined for each link to determine whether there is a problem (if the error count for any particular link exceeds the lower threshold, then there is a problem).

[0038] If, for any of the above methods, a timestamp is provided along with the count, endpoint 313 can additionally consider, for example, whether the link error is associated with any of the currently lost packets of the stream. For example, if receiving endpoint 313 is tracking a steady influx of telemetry information and detects a sudden spike in link errors within the same time window in which expected packets failed to arrive, receiving endpoint 313 can assume that its packets are among those included in the spike in errors. In this case, endpoint 313 can determine that there is a problem with stream 311 and, for example, generate a flag that causes endpoint 313 to request that sending endpoint 312 retransmit them or send an alert message to network management system 305.

[0039] For any of these methods, if a link ID is provided along with the link's telemetry, endpoint 313 can not only determine that there is an error in its flow, but can also name the link in the flow and / or the specific link in the flow that may be the source of the problem. The endpoint can send this information, for example, within an alert message sent to network management system 305. Such information can simplify the network management system's recovery process (e.g., by reconfiguring the switching table to avoid using the bad link).

[0040] Note that the destination endpoint 313 may collect telemetry and process it to make decisions / determinations and generate indicia in response thereto, or simply collect telemetry and send it to the source endpoint 312, which processes it to make decisions / determinations and generate indicia in response thereto. Operating points between these two extremes are also possible, where both endpoints 312, 313 perform some processing of the telemetry data and / or make decisions thereon.

[0041] Where the destination endpoint 313 sends telemetry information back to the source endpoint 312, the source endpoint may use the telemetry information, for example, to adjust one or more of the transmission parameters of the flow (eg, packet transmission rate, packet size, etc.).

[0042] Other possible methods of collecting telemetry and their subsequent processes are provided in 1) through 3) immediately below.

[0043] 1) Where the telemetry information includes a timestamp for each node hop traversed by one or more packets belonging to the flow, the destination endpoint 313 receiving the timestamp telemetry can construct a record of the end-to-end propagation delay (for a single packet) or the average end-to-end propagation delay (for multiple packets) through the network. The source endpoint 312 and / or the destination endpoint 313 can use this information, along with link quality telemetry and / or packet loss / corruption indicators (e.g., as per Figure 3a The above-mentioned alert message (such as the one about the packet loss detection timeout window) generates a flag that causes the endpoint protocols of the flow to tighten their packet loss detection timers (e.g., reduce the packet loss detection timeout window). These timeouts are usually set very conservatively (extended in time) to prevent false detections. In the case of bad link telemetry and / or increased flow rate related alert messages (such as the one about the packet loss detection timeout window above), the packet loss detection timeout window is increased. Figure 3a Those described), propagation delay telemetry information can be used to establish more sensible timeout windows (e.g., some moderate extension in time outside the core distribution of flows experiencing propagation delay), so that real errors are caught more quickly than with long timeout windows. Even if there is no indication of poor link quality in the telemetry, propagation delay information can be used to guide the setting of the timeout window.

[0044] 2) When source endpoint 312 and / or destination endpoint 313 learn that packet loss is more likely along the path of a particular flow and / or along a particular link, combined with the lack of any telemetry information indicating congestion within the network, they can determine that the link is experiencing, for example, noise or other deeper issues unrelated to the link's load (the link is bad). In this case, rather than generating a flag that causes, for example, a reduction in the sending rate or a decrease in the congestion window, the endpoint can generate a flag that causes the flow to be rerouted to avoid the link. This is an improvement over protocols that assume excessive sending rates and / or congestion cause packet loss. Instead, the endpoint generates a flag indicating "poor link quality causing packet loss," which makes no attempt to adapt the sending rate or reduce the congestion window for the affected flow.

[0045] 3) If source endpoint 312 and / or destination endpoint 313 learns of poor link quality along a flow used to send very small messages (e.g., messages consisting of only one or two packets), or the last packet of a message along that flow, this information can be used to cause source endpoint 312 to send the packet twice. Here, if the message consists of only one packet, for example, the loss of this packet will go unnoticed until the packet loss timer expires (because no subsequent packets in the flow can convey information about the lost packet). This can be time-consuming and lead to significant performance degradation. If the probability of packet loss is high, sending the message more than once (e.g., twice) increases the likelihood that at least one of them will reach the destination, thereby avoiding timeout penalties. At the same time, the bandwidth overhead of multiple transmissions is not significant, particularly if the packets are small. Multiple transmissions can also be used for very important messages that are, for example, time-sensitive or otherwise sensitive to the loss of any particular packet within the message's ordered packet stream.

[0046] It is worth noting that Figure 3a and Figure 3b The above-described method describes a rapid understanding by the endpoints 312, 313 of a flow of the performance of network devices within the network 304 supporting the flow 311. Here, the end-to-end propagation delay along the flow 311 through the network 304 can be less than 10 microseconds. Thus, even if a problem is not discovered, e.g., until the packet reaches its destination 313, the problem can still be discovered, diagnosed (e.g., using data that isolates the location within the network 304 where the problem occurred), and flagged to result in corrective action, long before traditional polling and analysis processor solutions would discover and report the problem.

[0047] Figure 3a In contrast, the method of is designed to immediately report problems with a particular packet via an alert message and a multi-bit error code from the switch that first drops the packet. Figure 3bThis process accumulates link telemetry across flows at the flow's destination endpoint 313. It's worth noting that a single link potentially supports a large number of different flows at any given moment. Therefore, specifically with respect to link telemetry, destination endpoint 313 is observing telemetry that affects / describes all flows flowing through the link shared by flow 311 and other flows. Therefore, it's conceivable that if a link experiences performance issues, the telemetry information received by multiple endpoints (for flows flowing through that link) could reveal the problematic link. Multiple flags generated simultaneously by multiple endpoints can further highlight the issue to network management.

[0048] Just now Figure 3b The discussion emphasizes a "feed-forward" approach, where telemetry information is collected on the ingress side and accumulated in the forward direction toward the receiving endpoint 313. In other embodiments, link telemetry information can be collected at the egress-side MAC layer 227 (e.g., and passed to the egress-side packet processing pipeline 225) for inclusion in egress packets (and accumulated with link telemetry for earlier links that the egress packet has traversed). As discussed above, the accumulated information can then be received at the receiving endpoint 313.

[0049] In yet other approaches, regardless of whether telemetry information is collected on the ingress and / or egress side of the switch, the telemetry information can be attached, alternatively or in combination, to packets being sent from destination endpoint 313 to source endpoint 312 (in the opposite flow direction). Sending telemetry to the source endpoint 312 of a flow allows source endpoint 312 to immediately generate a flag and take responsive corrective action, where source activity can mitigate the issue that generated the flag. For example, the source can begin retransmitting packets for any packets sent shortly after detecting an error spike along the flow.

[0050] In yet another method, reference Figure 3c , affecting per-flow telemetry. Here, telemetry information for a single flow is appended to the packets belonging to that flow. Problems in a flow can be detected at one of the flow's node hops (switches) or at one or both of the flow's endpoints.

[0051] Here, when the MAC layer of any of switches A, B, or C determines that it has received a damaged packet, it not only forwards the presence of an error (and possibly additional information, such as an error code specifying the type of error) to the packet processing pipeline within that switch, but also forwards the packet's source and destination address information and other header information to the pipeline (if the source and destination address information is deemed valid).

[0052] In this case, the packet processing pipeline can use this information to build tables that are populated with error statistics based on source and destination address information and / or other header information used to define the flow. Thus, telemetry information is collected on a per-flow basis. The telemetry information for each flow is then included in the packets belonging to that flow (e.g., within the header or as an additional payload). Telemetry can be the same as described above with respect to Figure 3a and Figure 3b Any of the telemetry in those described differs only in that they are specific to a particular flow, rather than reflecting the accumulation of all flows along a particular link.

[0053] Here, we can combine the above Figure 3a and Figure 3b The described mechanism can be used to detect problems in specific flows. For example, thresholds for errors or error rates can be predetermined and programmed into switches A, B, and C. If any of switches A, B, and C detects that the internal error count / error rate of a flow exceeds its threshold, the switch can send an alarm message to either or both of the flow's endpoints 312, 313 and / or the network management system 305. The alarm message includes a multi-bit error code that, for example, indicates that a threshold has been exceeded for a particular error count / error rate. The circuit module that collects telemetry for each flow and compares it to one or more applicable thresholds, as well as the circuit module that constructs the alarm message, can be integrated into the MAC layer circuit module and / or either or both of the ingress packet processing pipeline and the egress packet processing pipeline. The alarm message can also be sent to the network management system.

[0054] Back to reference Figure 2 , Figure 2 Shows different areas of the circuit module, which can achieve the above Figure 3a 、 Figure 3b and Figure 3c Alarm message generation and telemetry collection and reporting mechanisms supported by any / all packet processing pipelines described. Here, MAC layer circuitry module 222 includes circuitry module 241 to pass any of the following information to packet processing pipelines 223, 225: 1) headers of damaged packets (including information indicating whether the headers are valid or invalid); 2) link telemetry information (e.g., link rate, packets / second, link error count, link error rate, etc.). As further described below, given each packet's header information (which contains information defining the packet's flow) and any impairments identified by the MAC layer circuitry module and / or pipeline for packets of that flow, the packet processing pipeline can collect telemetry information for each flow.

[0055] This information can be forwarded directly from the ingress-side MAC layer circuitry module 222 to the ingress-side pipeline 223. According to a first approach, the information is forwarded to the pipeline 223 as a discrete data item. According to a second approach, the information is "piggybacked" using a valid packet passed from the MAC layer 222 to the pipeline 223. According to a third approach, the MAC layer 222 constructs a special packet with the information (e.g., with the information in its payload) and forwards the specially constructed packet to the ingress pipeline 223.

[0056] To pass information from the ingress MAC layer 222 to the egress pipeline 225, the ingress MAC layer 222 or the ingress pipeline 223 can construct a special packet that identifies, by destination address, where any alarm messages or telemetry submission reports generated from the information are to be sent. Alternatively, the information can be appended to a valid packet with the destination address. The packet is then switched by the switching core 204 and directed to the appropriate egress packet processing pipeline 225. The information is then processed by the egress pipeline 225, and any alarm messages and / or telemetry submission reports are generated as appropriate.

[0057] In the case of 1) above (a damaged packet header is forwarded to the pipeline), circuit module 241, along with MAC layer 222, is designed to forward the packet header to packet processing pipelines 223 and 225 if MAC layer 222 determines that the packet is damaged. (Circuit module 241 may also include information indicating whether the header is valid.) Therefore, even if the error checking circuit module within MAC circuit module 222 determines that, for example, the packet's payload is damaged after processing parity, cyclic redundancy check (CRC), error correction coding (ECC), forward error correction (FEC), or other error checking information included with the packet, circuit module 241 will still pass the packet header to packet processing pipelines 223 and 225.

[0058] The packet processing pipelines 223, 225 include circuit modules 242 to (as described above with respect to Figure 3a The invention relates to a method for constructing an alert message packet with a multi-bit error code using valid packet header information of a damaged packet, the alert message packet including the valid packet header in its payload and including a destination address in its header, the destination address being sufficient to send the alert message packet to a sending endpoint of a flow of damaged packets, a receiving endpoint of a flow of damaged packets, and / or a network management system.

[0059] In the event that the packet header is invalid, the circuit module 242 within the ingress packet processing pipeline 223 will append the damaged packet header to, for example, at least one valid packet of each flow being processed by the ingress pipeline 223, so that the damaged packet header will be received at the source or destination endpoint of each flow currently being processed by the pipeline 223.

[0060] Here, the ingress packet processing pipeline 223 includes a stage that performs packet classification. To perform packet classification, the stage maintains a table (eg, in a memory coupled to the stage) that has, for example, a separate entry for each flow that the pipeline 223 currently supports.

[0061] Here, pipeline circuit module 242 can maintain information for each entry indicating whether the pipeline has already attached a specific invalid header to a packet belonging to the flow of that entry. When pipeline 223 processes each new packet, pipeline 223 searches the entry for the flow for that packet for this information. If the entry indicates that an invalid header has been attached to a previous packet belonging to that flow, pipeline 223 does not attach an invalid header to the packet. If the entry indicates that an invalid header has not been attached to any previous packet belonging to that flow, pipeline 223 attaches the invalid header to the packet and updates the entry to indicate that an invalid header has been attached to a packet belonging to that flow.

[0062] Regarding the transmission of link telemetry, the MAC layer circuit module 241 can report the above information to the ingress pipeline 223. Figure 3b and Figure 3c Error information for any / all links and / or each flow described. This includes link rate, packets / second, link error count, link error rate, total link errors since reset, total link errors within the time window, and errors per flow (where source / destination address information (if valid) for corrupted packets is passed to the pipeline), etc.

[0063] Circuit module 242 of ingress packet processing pipeline 223 can also inject telemetry information into the header information of the packets it processes and / or create new packets containing telemetry information and appropriate header information (such as the correct source / destination addresses). Here, circuit module 242 is coupled to a memory that stores telemetry information. When telemetry information is passed from MAC layer 222 to pipeline 223, circuit module 242 writes the telemetry information to the memory. While pipeline 223 is processing a packet, circuit module 242 reads the telemetry information from a table and incorporates / injects the telemetry information into the packet. Alternatively, or in combination, circuit module 242 can create a new packet containing telemetry information and inject it into the stream supported by pipeline 223. This injection can be, for example, periodic, event-based, or the like. In the case of link telemetry (as opposed to per-stream telemetry), the MAC layer can instead include a circuit module coupled to a memory that stores the link telemetry information and performs any / all of these functions.

[0064] In the case of per-stream telemetry, the telemetry information in memory is viewed as a table with a different entry for each stream supported by pipeline 223. In this case, pipeline circuitry module 242 writes telemetry for a particular stream (e.g., packet error counts / error rates for various types of errors observed in packets of that stream, etc.) into the corresponding entry in the table for that stream and injects such telemetry only into packets belonging to that stream. Alternatively, or in combination, pipeline circuitry module 242 can create a new packet for a particular stream containing the telemetry information for that stream and inject it into that particular stream. Injection can be, for example, periodic, event-based, or the like.

[0065] Figure 4 Another feature of reporting errors and / or telemetry through the packet processing pipeline of the networking hardware switch 420 is shown. Specifically, as explained in the above teachings, damaged packets have been assumed to be damaged before being received by the ingress MAC layer of the switch. It is also possible that the switch 420 itself may have damaged the packet.

[0066] Here, Figure 4 An exemplary path 451 of a packet is shown. The packet is received losslessly along one of the ingress links 421 and processed without issue by the subsequent MAC layer circuitry 422 and ingress packet processing pipeline 423 (the packet remains lossless). The packet is then passed through a queue (not shown) and switched by the switch core 424. While the packet is traversing the queue and switch core 424, the packet's payload becomes corrupted. The egress packet processing pipeline 425, which processes the packet in the outbound direction (e.g., by processing error code information associated with the payload), detects the corruption.

[0067] Here, rather than allowing the packet to propagate along the egress link, packet processing pipeline 425 reroutes the packet back to ingress MAC layer circuitry module 422, which originally processed it after receiving it. The MAC layer circuitry module then continues to process the packet according to any of the above-described procedures, as if it had been received as a damaged packet, with error information passed to one of packet processing pipelines 423 or 425. However, it is important to note that the error information is tracked / recorded as being associated with the internal damage of switch 420 (rather than the link). Therefore, an additional dimension to the error statistics can specify whether the error is a link error or an internal switch error.

[0068] Note that while the above teachings have been directed to a networking switch 220 having a switching core 224, in various embodiments the switching core 224 is implemented using a routing core that converts ingress traffic to egress traffic through the execution of software executing on one or more processors rather than dedicated hardware circuit modules.

[0069] The various embodiments described above encompass implementations in which the packet processing pipeline incorporates error information received from the MAC layer circuit module (e.g., as is) into an alert message or other packet for error recovery, as well as implementations in which the packet processing pipeline processes the error information in some manner. For example, in the latter case, the packet processing pipeline accumulates previous error statistics with its local error statistics (see, e.g., Figure 3b and its discussion), and / or, the packet processing pipeline determines error statistics for each flow (see e.g. Figure 3c and its discussion).

[0070] Therefore, in addition to copying the "first" error information as received from the MAC layer into the packet used for error recovery (e.g., the alert message packet), the packet processing pipeline may also (or alternatively) calculate / determine "second" error information from such "first" error information and incorporate the second error information into the packet used for error recovery.

[0071] Any / all of the stream source and stream destination endpoint processes described above can be implemented using a stream source endpoint processing circuit module and a stream destination endpoint circuit module, respectively. Such circuit modules can be implemented using dedicated hardwired circuit modules (e.g., ASICs), programmable circuit modules (e.g., FPGAs), circuit modules that execute program code (e.g., processors), or any combination thereof.

[0072] Various aspects of the above teachings may be implemented to conform to various industry standards or specifications, such as the "In-Band Network Telemetry (INT) Data Plane Specification" v2.1 or later released by the P4.org Application Working Group on November 11, 2020.

[0073] about Figure 1a A new paradigm for computing environments (e.g., data centers) is emerging in which “infrastructure” tasks are offloaded from traditional general-purpose “host” CPUs (where application software programs are executed) to infrastructure processing units (IPUs), data processing units (DPUs), or smart networking interface cards (SmartNICs), any / all of which are hereinafter referred to as IPUs.

[0074] Network-based computer services (such as those provided by cloud services and / or large enterprise data centers) typically execute application software programs on behalf of remote clients. In this context, the application software programs typically perform specific (e.g., "business") end functions (e.g., customer service, procurement, supply chain management, email, etc.). Remote clients invoke / use these applications via a temporary network session / connection established between the client and the application by the data center.

[0075] However, to support the functionality of the network session and / or application, certain underlying computationally intensive and / or traffic intensive functions ("infrastructure" functions) are performed.

[0076] Examples of infrastructure functionality include encryption / decryption for secure network connections, compression / decompression for small footprint data storage and / or network communications, virtual networking between clients and applications and / or between applications, packet processing, ingress / egress queuing of networking traffic between clients and applications and / or between applications, ingress / egress queuing of command / response traffic between applications and mass storage devices, error checking (including checksum calculations to ensure data integrity), distributed computing remote memory access capabilities, etc.

[0077] Traditionally, these infrastructure functions have been performed by CPU units "beneath" their end-function applications. However, the density of infrastructure functions has begun to impact the CPU's ability to execute their end-function applications in a timely manner relative to client expectations and / or in an energy-efficient manner relative to data center operators' expectations. Furthermore, processes executing a wide variety of different application software programs make better use of CPUs, which are typically complex instruction set (CISC) processors, than more monotonous and / or focused infrastructure processes.

[0078] Therefore, if Figure 5 As observed in [1], infrastructure functions are being moved to infrastructure processing units. Figure 5 An exemplary data center environment 500 is depicted that integrates an IPU 507 to offload infrastructure functions from a host CPU 504, as described above.

[0079] like Figure 5 As viewed in FIG, an exemplary data center environment 500 includes a pool of CPU units 501 that execute final function application software programs 505 typically invoked by remote call clients. The data center also includes separate memory pools 502 and mass storage device pools 405 to assist in executing applications.

[0080] The CPUs, memory storage, and mass storage pools 501, 502, 503 are each coupled via one or more networks 504. The network(s) may include switches and / or routers that use packet processing pipelines to (as discussed above with respect to Figure 2 、 Figure 3a 、 Figure 3b 、 Figure 3c and Figure 4 Detailed description of the .NET Framework 5.1.1.1.1 and .NET Framework 5.1.1.1.1 (see .NET Framework 5.1.1.1 for details) to track, report, and recover from network errors.

[0081] Notably, each pool 501, 502, and 503 has an IPU 507_1, 507_2, or 507_3 on its front end, or network side. Each IPU 507 performs pre-configured infrastructure functions on incoming (request) packets it receives from network 504 before delivering these requests to its corresponding pool's final function (e.g., executing software in the case of CPU pool 501, memory in the case of memory pool 502, and storage in the case of mass storage pool 503). When the final function sends certain communications onto network 504, the IPU 507 performs the pre-configured infrastructure functions on the outbound communications before passing them onto network 504.

[0082] Depending on the implementation, one or more CPU pools 501, memory pools 502, and mass storage pools 503, along with network 504, may reside within a single chassis, for example, as a traditional rack-mounted computing system (e.g., a server computer). In a disaggregated computing system implementation, one or more CPU pools 501, memory pools 502, and mass storage pools 503 are separate rack-mountable units (e.g., rack-mountable CPU units, rack-mountable memory units (M), rack-mountable mass storage units (S)).

[0083] In various embodiments, the software platform on which the application 505 is executed includes a virtual machine monitor (VMM) or hypervisor that instantiates multiple virtual machines (VMs). Operating system (OS) instances are respectively executed on the VMs, and applications are executed on the OS instances. Alternatively or in combination, container engines (e.g., Kubernetes container engines) are respectively executed on the OS instances. The container engine provides a virtualized OS instance, and the containers are respectively executed on the virtualized OS instances. The container provides an isolated execution environment for a suite of applications, which may include applications for microservices. The same software platform can be used in Figure 2 is executed on the CPU unit 201.

[0084] Figure 6a An exemplary IPU 607 is shown. As seen in FIG6 , the IPU 609 includes a plurality of general-purpose processing cores 611, one or more field programmable gate arrays (FPGAs) 612, and / or one or more acceleration hardware (ASIC) blocks 613. The IPU typically has at least one associated machine-readable medium to store software to be executed on the processing cores 611 and firmware to program the FPGAs (if present) so that the processing cores 611 and FPGAs 612 (if present) can perform their intended functions.

[0085] Processing core 611, FPGA 612, and ASIC block 613 represent different trade-offs between versatility / programmability, computational performance, and power consumption. Generally speaking, tasks can be performed faster and with minimal power consumption in an ASIC block; however, an ASIC block is a fixed-function unit that can only perform the function for which its electronic circuit module has been specifically designed.

[0086] In contrast, the general-purpose processing cores 611 will perform their tasks more slowly and consume more power, but can be programmed to perform a wide variety of different functions (via the execution of software programs). It is worth noting that while the processing cores can be general-purpose CPUs, such as the host CPU 501 in a data center, in many instances the general-purpose processors 511 of the IPU are reduced instruction set computing (RISC) processors, rather than CISC processors (which are typically implemented as CISC processors in the host CPU 501). In other words, the host CPU 501 that executes the data center's application software programs 505 is often a CISC-based processor, because the data center's application software may be programmed to perform a wide variety of different tasks (regarding the host CPU 501). Figure 2 , the CPU unit 201 is typically also a general-purpose CISC processor).

[0087] In contrast, the infrastructure functions performed by the IPU tend to be a more limited set of functions that are more suited to being provided by a RISC processor. Therefore, the RISC processor 611 of the IPU should perform infrastructure functions with less power consumption than a CISC processor without a significant loss of performance.

[0088] FPGA(s) 612 provide more programming capability than an ASIC block, but less programming capability than a general purpose core 611 , while at the same time providing more processing execution capability than a general purpose core 611 , but less processing execution capability than an ASIC block.

[0089] Figure 6b A more specific embodiment of the IPU 607 is shown. Figure 6b The specific IPU 607 does not include any FPGA blocks. Figure 6bAs observed in , the IPU 607 includes multiple general-purpose cores (eg, RISC) 611 and a last-level cache layer for the general-purpose cores 611 . The IPU 607 also includes a plurality of hardware ASIC acceleration blocks, including: 1) an RDMA acceleration ASIC block 621 that performs RDMA protocol operations in hardware; 2) an NVMe acceleration ASIC block 622 that performs NVMe protocol operations in hardware; 3) a packet processing pipeline ASIC block 623 that parses the ingress packet header content, for example, to assign a flow to the ingress packet, perform network address translation, etc.; 4) a traffic shaper 624 that assigns the ingress packet to the appropriate queue for subsequent processing by the IPU 509; 5) an inline cryptographic ASIC block 625 that performs decryption on the ingress packet and encryption on the egress packet; 6) a bypass cryptographic ASIC block 626 that performs encryption / decryption on a data block, for example, as requested by the host CPU 501; 7) a bypass compression ASIC block 627 that performs compression / decompression on a data block, for example, as requested by the host CPU 501; 8) a checksum / cyclic redundancy check (CRC) calculation (for example, for NVMe / TCP data digest and / or NVMe DIF / DIX data integrity); 9) Thread Local Storage (TLS) procedures; etc.

[0090] The packet processing pipeline 623 may be included at any one or more of the constituent stages of the pipeline (as discussed above with respect to Figure 2 、 Figure 3a 、 Figure 3b 、 Figure 3c and Figure 4 Detailed description of the IPU 607 includes functionality for tracking, reporting, and recovering from network errors. It is conceivable that the IPU 607 can be considered a network component (e.g., at the edge of the network 504).

[0091] The IPU 507 also includes multiple memory channel interfaces 628 coupled to external memory 629 for storing instructions for the general purpose core 511, as well as input / output data for each of the IPU core 511 and the ASIC blocks 621-626. The IPU includes multiple PCIe physical interfaces and an Ethernet media access control block 630 to enable network connectivity to and from the IPU 609. As mentioned above, the IPU 607 can be a semiconductor chip or multiple semiconductor chips integrated on a module or card (e.g., a NIC).

[0092] Figure 7 An embodiment of an ingress packet processing pipeline 703 is shown, which may include any one or more of the constituent stages of the pipeline (as described above with respect to Figure 2 、 Figure 3a 、 Figure 3b 、 Figure 3c and Figure 4 Circuit modules that track and report network errors and / or telemetry (e.g., Figure 2 circuit module 242). Figure 7 As observed in , the packet processing pipeline 703 is used to process inbound packets and assign each inbound packet to the appropriate queue. Generally speaking, the pipeline 703 includes a stage 704 located at (or towards) the front end of the pipeline that parses the header of the packet and extracts information found in the various fields of the header.

[0093] Pipeline 703 also includes another stage 705 that identifies the flow to which the inbound packet belongs, or otherwise "classifies" the packet for downstream processing or handling ("packet classification"). Here, the extracted packet header information (or portion(s) thereof) is compared with entries in a table of lookup values ​​708. The particular entry whose value matches the packet's header information identifies the flow to which the packet belongs, or otherwise classifies the packet.

[0094] The packet processing pipeline 703 also includes a stage 706 located at (or towards) the back end of the pipeline, which directs the packet to a specific one of the inbound queues 702_1 to 702_N based on the content of the header information of the inbound packet (typically the port and IP address information of the source and destination of the packet).

[0095] Typically, packets with the same source and destination header information belong to the same flow and will be assigned to the same queue. Because each queue is associated with a specific quality of service (e.g., queue service rate), switch core input port, or other processing core, forwarding inbound packets with the same source and destination information to the same queue enables common processing of packets belonging to the same flow.

[0096] The egress pipeline may also be multi-stage and may be used to prepare packets for outgoing transmission (eg, at Layer 3 (IP layer) or higher), such as creating IP header information for outbound packets.

[0097] Embodiments of the present invention may include various processes as described above. These processes may be implemented using program code (e.g., machine-executable instructions). When the program code is processed, it causes a general-purpose or special-purpose processor to execute the process of the program code. Alternatively, these processes may be performed by dedicated / custom hardware components comprising hardwired interconnected logic circuit modules (e.g., application-specific integrated circuit (ASIC) logic circuit modules) or programmable logic circuit modules (e.g., field-programmable gate array (FPGA) logic circuit modules, programmable logic device (PLD) logic circuit modules) for performing the processes; or they may be performed by any combination of program code and logic circuit modules.

[0098] Elements of the present invention can also be provided as machine-readable media for storing program code.Machine-readable media can include, but are not limited to, floppy disks, optical disks, CD-ROMs and magneto-optical disks, FLASH memories, ROMs, RAMs, EPROMs, EEPROMs, magnetic cards or optical cards or other types of media / machine-readable media suitable for storing electronic instructions.

[0099] In the foregoing description, the present invention has been described with reference to specific exemplary embodiments thereof. However, it will be apparent that various modifications and variations may be made thereto without departing from the broader spirit and scope of the present invention as set forth in the appended claims. Accordingly, the description and drawings are to be regarded as illustrative rather than restrictive.

Claims

1. A device comprising: An electronic circuit module is configured to support a plurality of flows within a network, the electronic circuit module being configured to determine corresponding telemetry information for the plurality of flows and, when an alarm condition is reached for a particular one of the plurality of flows, inject an alarm message into the particular one of the plurality of flows, the alarm message including a multi-bit error code describing the alarm condition, the multi-bit error code being one of a plurality of possible multi-bit error codes.

2. The device according to claim 1, wherein The electronic circuit module includes a packet processing pipeline.

3. The device according to claim 1 or 2, wherein: The packet processing pipeline is coupled between the ingress media access control circuit module and the switch core.

4. The device according to claim 1 or 3, wherein The electronic circuit module is used to insert the alarm message into the header of the packet belonging to the specific one flow.

5. The device according to claim 1 or 3, wherein: The electronic circuit module is used to create a packet carrying the alarm message and inject the packet into the specific one flow.

6. The device according to claim 1 or 3, wherein: The electronic circuit module is used for collecting corresponding telemetry data for the multiple streams and injecting the corresponding telemetry data into the multiple streams respectively.

7. The apparatus according to claim 6, wherein The electronic circuit module is to accumulate the corresponding telemetry data for a specific other flow of the plurality of flows with earlier telemetry data that was earlier determined for the specific other flow of the plurality of flows and received at the ingress media access control interface.

8. The apparatus according to claim 7, wherein The electronic circuit module is used to inject the accumulated telemetry data into the specific other stream among the multiple streams.

9. The apparatus according to claim 6, wherein The corresponding telemetry data for the particular another one of the plurality of streams includes telemetry data for a link that transmits packets of the particular another one of the plurality of streams and a plurality of other ones of the plurality of streams.

10. A device comprising: a stream endpoint processing circuit module for processing packets of the stream at an endpoint of the stream, the stream endpoint processing circuit module for processing packets belonging to the stream, the packets including an alarm message, the alarm message including a multi-bit error code describing an alarm condition that has been reached for the stream, the multi-bit error code being one of a plurality of possible multi-bit error codes, the stream endpoint processing circuit module for generating a flag in response to the processing of the packet by the stream endpoint processing circuit module.

11. The apparatus according to claim 10, wherein The flow endpoint processing circuit module is a flow destination endpoint processing circuit module, and the flow destination endpoint processing circuit module is used to alert the source endpoint of the flow to the alarm message in response to the flag.

12. The apparatus according to claim 10, wherein The flow endpoint processing circuit module is a flow destination endpoint processing circuit module, and the flow destination endpoint processing circuit module is configured to send a second alarm message to a network management function in response to the flag.

13. A data center, comprising: CPU pool; Memory resource pool; accelerator pool; A network communicatively coupling the CPU pool, the memory resource pool, and the accelerator pool, the network comprising a network switch, the network switch comprising the following a), b), c), and d): a) Switch core; b) Ingress media access control circuit module; c) an egress media access control circuit module, the network switch being configured to support a plurality of flows flowing from the egress media access control circuit module through the switch core into the ingress media access control circuit module; as well as d) an electronic circuit module for supporting the plurality of streams, the electronic circuit module being configured to determine corresponding telemetry data for the plurality of streams and, when an alarm condition is reached for a particular one of the plurality of streams, inject an alarm message into the particular one of the plurality of streams, the alarm message including a multi-bit error code describing the alarm condition, the multi-bit error code being one of a plurality of possible multi-bit error codes.

14. The data center according to claim 13, wherein: The electronic circuit module includes a packet processing pipeline coupled between the ingress media access control circuit module and the switch core.

15. The data center according to claim 13 or 14, wherein: The electronic circuit module is used to insert the alarm message into the header of the packet belonging to the specific one flow.

16. The data center according to claim 13 or 14, wherein: The electronic circuit module is used to create a packet carrying the alarm message and inject the packet into the specific one flow.

17. The data center according to claim 16 further includes a stream endpoint processing circuit module to process packets of the specific one stream at the endpoint of the specific one stream, the stream endpoint processing circuit module is used to process packets belonging to the stream and containing the alarm message, and the stream endpoint processing circuit module is used to generate a tag in response to the processing of the packet by the stream endpoint processing circuit module.

18. The data center according to claim 17, wherein: The endpoint is a source endpoint.

19. The data center according to claim 17, wherein: The endpoint is the destination endpoint.

20. The data center of claim 17, wherein: The circuit module includes a packet processing pipeline.