Pass-through delay and network fault diagnosis with limited error propagation

By introducing an automatic switching mechanism between CT and SAF modes in network switches, the contradiction between latency and fault isolation in existing technologies is resolved, enabling fault isolation and data integrity verification under low latency, thereby improving the efficiency and security of network management.

CN116805934BActive Publication Date: 2025-11-21AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310306378.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-25
Filing Date
2023-03-24
Publication Date
2025-11-21
Estimated Expiration
2043-03-24

AI Technical Summary

Technical Problem

Existing network switches cannot simultaneously achieve the dual goals of low latency and fault isolation. CT switches have poor fault isolation issues, while SAF switches cause significant latency.

Method used

Design a switch with CT mode and SAF mode, which automatically switches operating modes based on port health indicators. When the health indicator drops below a threshold, switch to SAF mode to verify data integrity and automatically alert the remote system before reverting to CT mode.

Benefits of technology

It enables effective isolation of faulty data packets without affecting latency, reduces network latency, simplifies fault diagnosis, prevents data leakage, and improves network management efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116805934B_ABST
    Figure CN116805934B_ABST
Patent Text Reader

Abstract

This application relates to cut-through delay and network fault diagnosis with limited error propagation. A switch can operate in a cut-through mode as well as a store-and-forward mode. When in the default cut-through mode, the switch continuously monitors certain health indicators of the ports. If those health indicators fall below a threshold, the switch changes to operate in the store-and-forward mode for a predetermined period of time or until the health indicators rise above the threshold, at which point the switch can resume cut-through mode operation. If the health indicators fall below an even lower threshold or remain below the threshold for a predefined period of time, the switch can automatically alert a remote system or software process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the inventive concepts disclosed herein generally relate to network switches, and more specifically to network switches that enter and leave pass-through mode based on port-specific health indicators. Background Technology

[0002] Existing network switches (such as those used in data centers) cannot meet the dual goals of low latency and fault isolation. Cut-through (CT) switching offers low latency but poor fault isolation because data packets may be passed through before the entire packet is received and therefore before a fault is identified. Store-and-forward (SAF) switching is fault-tolerant because it requires receiving the entire packet before forwarding, but it introduces significant latency into packets that would otherwise be safely passed through. Switches and switching methods that achieve both low latency and fault isolation would be advantageous. Summary of the Invention

[0003] In one aspect, embodiments of the inventive concept disclosed herein relate to a switch having CT mode and SAF mode. In the default CT mode, the switch continuously monitors certain health metrics of its ports. If those health metrics drop below a threshold, the switch switches to SAF mode for a predetermined period of time or until the health metrics rise above the health threshold, at which point the switch can revert to CT mode operation.

[0004] Furthermore, if a health indicator drops to or remains below a threshold or even lower for a predefined period of time, the switch can automatically alert a remote system or software process.

[0005] It should be understood that the foregoing overview and the following detailed description are merely exemplary and explanatory and should not limit the scope of the claims. The accompanying drawings, which are incorporated in and form part of this specification, illustrate exemplary embodiments of the inventive concepts disclosed herein and, together with the overview, serve to explain the principles. Attached Figure Description

[0006] By referring to the accompanying drawings, those skilled in the art can better understand the numerous advantages of the embodiments of the inventive concepts disclosed herein, in which:

[0007] Figure 1A A diagram illustrating the propagation of data packets through a switch in different operating modes;

[0008] Figure 1B A diagram illustrating the propagation of data packets through a switch in different operating modes;

[0009] Figure 2AA box diagram illustrating the propagation of malicious data packets through a switch network in CT mode;

[0010] Figure 2B A block diagram illustrating the propagation of malicious data packets through a switch network in SAF mode;

[0011] Figure 3 A block diagram illustrating the propagation of data packets through a switch under various fault conditions;

[0012] Figure 4A A block diagram illustrating the propagation of data packets through a CT switch;

[0013] Figure 4B A block diagram illustrating the propagation of data packets through an SAF switch;

[0014] Figure 5 A block diagram illustrating the propagation of data packets through a switch with a secure boundary;

[0015] Figure 6 A block diagram showing a switch according to an exemplary embodiment of the present disclosure;

[0016] Figure 7 Histograms of identified faults useful in exemplary embodiments of this disclosure are shown;

[0017] Figure 8 A block diagram illustrating the propagation of data packets through a switched network according to an exemplary embodiment of the present disclosure;

[0018] Figure 9 A block diagram illustrating the propagation of data packets through a switched network according to an exemplary embodiment of the present disclosure;

[0019] and

[0020] Figure 10 A block diagram illustrating the propagation of data packets through a switched network according to an exemplary embodiment of the present disclosure. Detailed Implementation

[0021] Before explaining in detail the various embodiments of the inventive concept disclosed herein, it should be understood that the inventive concept is not limited to the arrangement of components, steps, or methods set forth in the following description or illustrated in the accompanying drawings. In the following detailed description of embodiments of the inventive concept, numerous specific details are set forth to provide a more thorough understanding of the inventive concept. However, those skilled in the art to which this disclosure pertains will understand that the inventive concept disclosed herein can be practiced without these specific details. In other instances, well-known features may not be described in detail to avoid unnecessarily complicating the disclosure. The inventive concept disclosed herein can have other embodiments or can be practiced or implemented in various ways. Furthermore, it should be understood that the wording and terminology used herein are for descriptive purposes and should not be considered limiting.

[0022] As used herein, the letters following the reference numerals are intended to refer to embodiments of features or elements that are similar to, but not necessarily identical to, previously described elements or features with the same reference numerals (e.g., 1, 1a, 1b). Such shorthand notation is for convenience only and should not be construed as limiting the inventive concepts disclosed herein in any way unless expressly stated otherwise.

[0023] Furthermore, unless explicitly stated otherwise, "or" refers to an inclusive "or" and not an exclusive "or". For example, condition A or B is satisfied by either of the following: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), and both A and B are true (or exist).

[0024] Additionally, the use of “a” or “an” is adopted to describe elements and components of embodiments of the inventive concept. This is done only for convenience and to give the general meaning of the inventive concept, and “a” and “an” are intended to include one or at least one, and the singular includes the plural, unless it is obvious that they have other meanings.

[0025] Furthermore, while various components may be depicted as directly connected, direct connection is not mandatory. Components may communicate data with intermediary components that are not specified or described. It is understood that "data communication" refers to both direct and indirect data communication (e.g., the possibility of intermediary components).

[0026] Ultimately, as used herein, any reference to “one embodiment” or “some embodiments” means that a particular element, feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of the inventive concept disclosed herein. The phrase “in at least one embodiment” appearing in the specification does not necessarily refer to the same embodiment. Embodiments of the disclosed inventive concept may include one or more features explicitly described herein or inherently present, or any combination or sub-combination of two or more such features.

[0027] In summary, embodiments of the inventive concepts disclosed herein relate to a network device having a CT mode and a SAF mode. In the default CT mode, the network device continuously monitors certain health indicators of its ports. If those health indicators drop below a threshold, the network device switches to SAF mode for a predetermined period of time or until the health indicators rise above the threshold, at which point the network device may revert to CT mode operation. If the health indicators drop below or even lower than the threshold, or remain below the threshold for a predefined period of time, the network device may automatically alert a remote system or software process.

[0028] refer to Figures 1A to 1B The diagram illustrates the propagation of data packets through a switch in different operating modes. In CT mode, switch 100 receives input data packet 104 and begins outputting output data packet 106 as early as possible before receiving the entire input data packet 104. In contrast, SAF mode switch 102 receives the entire input data packet 104 and verifies the integrity of the input data packet before outputting output data packet 106. Compared to SAF mode switch 102, CT mode switch 100 exhibits less latency 108, 110 between receiving input data packet 104 and outputting output data packet 106.

[0029] Although input data packet 104 is fault-free, CT mode switch 100 propagates data packets 104 and 106 faster than SAF mode switch 102. However, faulty data packets travel through CT mode switch 100 before the fault is identified. For example, faulty input data packet 112 (e.g., a data packet containing errors identified via cyclic redundancy check, error correction code, cryptographic hash function, or the like) may be passed through CT mode switch 100 as faulty output data packet 114 because the error check bit only appears at the end of data packets 112 and 114. Although both CT mode switch 100 and SAF mode switch 102 will identify faulty input data packet 112, only SAF mode switch 102 will stop propagating it.

[0030] By reducing switch latency, CT mode switch 100 has lower network latency; the benefit increases by a factor of hops, as each hop in SAF mode switch 102 adds incremental switch latency. However, when a faulty input data packet 112 is present, the network of CT mode switch 100 can propagate the faulty input data packet 112 without restriction.

[0031] refer to Figures 2A to 2BThis diagram illustrates the propagation of faulty data packets through a switch network in different modes. In the network of CT mode switches 200, 202, 204, 206, 208, and 210, faulty data packets 224, 226, 228, and 230 can be propagated to any other CT mode switch 200, 202, 204, 206, 208, or 210, and even to end users / clients. For example, the first CT mode switch 200 can receive the first faulty data packet 224 and the second faulty data packet 226 via the first port 232. The first faulty data packet 224 and the second faulty data packet 226 can be rapidly propagated to downstream CT mode switches 202, 204, 206, 208, and 210. At each hop (propagating from one switch 200, 202, 204, 206, 208, 210 to the next switch), switches 200, 202, 204, 206, 208, 210 may determine that faulty data packets 224, 226, 228, 230 are faulty based on some error checking mode, and mark the corresponding ports 232, 234, 236, 238 as supplying faulty data packets 224, 226, 228, 230, but will continue to propagate faulty data packets 224, 226, 228, 230. It should be understood that, in the context of this disclosure, "port" refers to a defined portion of the address space of a device in which data can be sent and received. Some examples of "port" or "data port" described herein may refer to a "receive port" or "relay port" or the like; in those examples, it will be understood that such distinctions are inherently relative and based on the instantaneous directionality of the data flow.

[0032] In contrast, in the network of SAF mode switches 212, 214, 216, 218, 220, and 222, faulty data packets 224, 226, 228, and 230 can still be received at ports 232 and 234 of the first SAF mode switch 212. However, the first SAF mode switch 212 will verify those faulty data packets 224, 226, 228, and 230 before retransmitting them, and will isolate the faulty data packets 224, 226, 228, and 230 at the first SAF mode switch 212. Further analysis can identify the faulty link or upstream device that provided the faulty data packets 224, 226, 228, and 230, thereby simplifying network diagnostics and management. However, although the data packets are valid, each SAF mode switch 212, 214, 216, 218, 220, and 222 imposes an additional delay per data packet.

[0033] refer to Figure 3This diagram illustrates the propagation of data packets through a switch under various fault conditions. The SAF mode switch 300 may experience several types of data packet failures. In the first scenario 302, a faulty upstream device generates some faulty data packets 308 and some valid data packets 310. The SAF mode switch will successfully propagate the valid data packets 310 and isolate the faulty data packets 308.

[0034] In the second scenario 304, the upstream device generates valid data packets 310, but connects to the SAF mode switch 300 via a faulty link 312 to a specific port. The faulty link 312 functionally causes all valid data packets 310 to appear faulty, and the SAF mode switch 300 will isolate all those valid data packets 310.

[0035] In scenario 306, a faulty upstream device generates some faulty data packets 308 and some valid data packets 310, but the upstream device is connected to the SAF mode switch 300 via a faulty link 312. The faulty link 312 functionally causes all valid data packets 310 to appear faulty, and the SAF mode switch 300 isolates both the faulty data packets 308 and the valid data packets 312.

[0036] It's clear that successfully isolating a fault to a specific source port is not a complete diagnosis. Network management must still determine whether the fault lies in an upstream device, a link, or both. The SAF mode switch 300 can log fault statistics for future planning and remediation.

[0037] refer to Figures 4A to 4BThe diagram illustrates the propagation of data packets through switches 400 and 406 in different modes. In the case where switch 400 in CT mode is connected to upstream servers 402 and 404 for downstream data propagation, from the perspective of switch 400 in CT mode, the first server 402 may generate valid data packets 408, 410, and 412, while the second server 404 may generate faulty data packets 414, 416, and 418 (due to a faulty data source or a faulty link). Although CT mode switches generally produce lower latency compared to SAF mode switches, in cases where CT mode switch 400 allows faulty data packets 414, 416, and 418 to pass through, these faulty data packets 414, 416, and 418 may be unavailable, thus resulting in a practically longer latency for valid data packets 408, 410, and 412 that propagate after them. In contrast, the SAF-mode switch 406 in the same scenario will generate lower network-level latency because the bandwidth within the network is not consumed by faulty data packets 414, 416, and 418. Discarding faulty data packets 414, 416, and 418 avoids the latency caused by valid data packets 408, 410, and 412 waiting after faulty data packets 414, 416, and 418. This consideration is especially important if the second upstream server 404 is sending an unusually high load compared to normal network traffic.

[0038] The failure rates and failure modes of servers (402, 404, and endpoints) may be significantly worse than those of upstream switches due to their higher overall complexity. For single-hop switching between CT and SAF modes, the impact on latency is negligible, but it has a significant impact on the system's ability to diagnose both network and server errors.

[0039] refer to Figure 5 This diagram illustrates the propagation of data packets through switches 500 and 502 with a security boundary 504. In a specific scenario, a faulty header in a faulty data packet 506 may lead to incorrect forwarding. In this scenario, the CT-mode switch 500 may incorrectly route the internal faulty data packet 506 as the external faulty data packet 508 based on the faulty header. However, the packet payload of the external faulty data packet 508 may still contain sensitive data that is forwarded outside the physical barrier.

[0040] refer to Figure 6This diagram illustrates a block diagram of a switch according to an exemplary embodiment of the present disclosure. The switch includes a controller 600 or processor configured to operate the switch in CT mode and SAF mode. It will be understood that "processor" can refer to a dedicated processor hardwired for the described purposes, a general-purpose programmable central processing unit (CPU), a field-programmable gate array (FPGA), and other such data processing technologies. Where the processor includes means configurable by software or firmware, this software or firmware may be embodied in a non-transitory memory; this memory may be a PROM, EPROM, EEPROM, flash memory, dynamic random access memory, or the like. The controller 600 is configured to perform certain process steps 606, 608, 610, 612, as described more fully herein.

[0041] By default, controller 600 maintains the switch in CT mode and continuously monitors the health of each connected data port. In CT mode, controller 600 is electronically configured to receive data packets from upstream receive port 602 and forward the data packets 606 to downstream trunk port 604. Controller 600 is configured to simultaneously perform a data integrity check 608 on each data packet or a sample of data packets to quantify port link health in the form of a port health metric. A port health metric is a quantification of known data integrity errors associated with the corresponding receive port 602. When the port health metric falls below a predefined unhealthy threshold, controller 600 is configured to change 610 to SAF mode. The predefined unhealthy threshold can be defined by the number of errors associated with receive port 602 over time, the ratio of faulty packets associated with receive port 602, a histogram offset as described herein, or the like. In SAF mode, controller 600 performs and completes data integrity checks (including any error correction, cyclic redundancy checks, cryptographic hash functions, and the like) on each data packet before relaying it to the corresponding downstream device. Faulty data packets are discarded when data correction is impossible, such as when the number of faulty bits is too large. Discarding faulty packets simplifies network diagnostics and management by isolating faulty components, avoids increased latency caused by waiting behind faulty data packets, protects the network from faulty servers and other endpoints, and prevents data leakage from physical security boundaries. In at least one embodiment, controller 600 is configured to change 612 back to CT mode after a predefined time period or after the port health metric of port 602 rises above a defined health threshold. The predefined health threshold may be defined by the number of errors over time associated with receive port 602, the ratio of faulty packets associated with receive port 602, histogram offset, or the like.

[0042] In at least one embodiment, the controller 600 relays data packets for an upstream port 602 in SAF mode where the port health metric is below a predefined unhealthy threshold, while continuing to operate the remaining upstream ports 602 in CT mode.

[0043] In at least one embodiment, port health metrics are measured by monitoring forward error correction (FEC) statistics via a flight data recorder. Similarly, upstream device health can be measured by monitoring an Ethernet MIB counter. Furthermore, cyclic redundancy check (CRC) can be used to identify faulty data packets and the degree of fault (number of faulty bits). Port health metrics can be measured by the number of faults per port 602 over a period of time, where the number of faults is a rolling counter. Alternatively or additionally, faults can be weighted using more recent faults that have a more significant impact on port health metrics than previous faults. Weighted fault measurements can identify trends in port health to allow controller 600 to proactively switch 610 to SAF mode before severe fault loads.

[0044] In at least one embodiment, controller 600 maintains a histogram or histogram-like graph of faults for each data port 602. In such embodiments, port health metrics may be at least partially measured as the offset of the histogram over time. With controller 600 registering the histogram offset over time, controller 600 may identify a trend of increasing fault bits and switch 610 to SAF mode before controller 600 is configured to relay any unrecoverable data packets 606. In at least one embodiment, controller 600 is configured to change 610 within 50 μs for 100G port 602.

[0045] When controller 600 has changed 610 to SAF mode, the CRC error counter will only continue to increment for the less healthy source port 602. The step of changing 610 to SAF mode limits the number of faulty data packets propagating to trunk port 604.

[0046] Temporarily operating in SAF mode allows controller 600 to communicate with the management plane monitor host processor 614 and alert the network health management system to faulty port 602. It will be understood that the management plane monitor may include any system for receiving and logging network events corresponding to errors and the like that may require some user intervention. If controller 600 determines that the health of port 602 has declined below a lower threshold defined by factors such as the number of errors over time, the ratio of faulty packets, histogram offset, or the like, then the controller may automatically notify host processor 614 and other software to share error statistics 616 that may be useful for diagnosing the fault. Embodiments of this disclosure allow for rapid alerting of the management plane monitor and error statistics 616 collection to guide control plane and device replacement decisions. In at least one embodiment, the switch may include more than one controller 600 to accelerate detection.

[0047] While the embodiments described herein specifically refer to a "switch," it should be understood that the embodiments are applicable to any computer device that receives and distributes data packets in a network and can operate in CT mode or SAF mode. In the context of this disclosure, "computer device" means any apparatus having one or more processors specifically configured to perform the functions described herein, or that can be electronically configured via software or firmware.

[0048] refer to Figure 7 The diagram illustrates a histogram of identified faults useful in exemplary embodiments of this disclosure. In at least one embodiment, the switch processor can track faults over time. For example, within a first time period 700, the processor can identify data packets with one fault 702, two faults 704, and three faults 706. The distribution of the number of faults 702, 704, and 706 in the faulty data packets during the first time period 700 can be represented as a first time period distribution curve 708. In a subsequent second time period 710, the processor can identify data packets with one fault 712, two faults 714, and three faults 716. The distribution of the number of faults 712, 714, and 716 in the faulty data packets during the second time period 710 can be represented as a second distribution curve 718.

[0049] The offset from the first time-period distribution curve 708 to the second time-period distribution curve 718, represented by basic data, can indicate port link health degradation as shown by the histogram offset corrected by FEC. For example, the offset from primarily one faulty 702 data packet to an increasing number of three faulty 716 data packets can indicate degraded data link quality. When an offset towards data packets with increasing faults can be identified in real time, the processor can switch to SAF mode before the port begins generating major unrecoverable faulty data packets, thereby preventing unrecoverable data packets from being forwarded downstream. Health metrics based on histogram offsets can identify failure trends before a large number of errors occur, thus reducing the overall impact on system performance.

[0050] In at least one embodiment, the number of faults 702, 704, 706, 712, 714, and 716 may represent corrected data symbols. When data packets comprise codewords consisting of data, symbols, and parity or check symbols, if a corrupted link exists, the receiver can perform FEC and correct the data symbols. By tracking the number of symbols requiring correction (e.g., represented by the number of faults 702, 704, 706, 712, 714, and 716) and the offset of the required symbol corrections over time, a faulty or degraded link can be identified before the symbols become uncorrectable by FEC.

[0051] refer to Figure 8This diagram illustrates the propagation of data packets through a network of switches 800, 802, 804, 806, 808, and 810 according to an exemplary embodiment of the present disclosure. When switches 800, 802, 804, 806, 808, and 810 are configured for CT mode or SAF mode, the first switch 800 can receive faulty data packets 824, 826, and 828 from a first port 832 and faulty data packet 830 from a second port 834. In the default CT mode, the first switch 800 can relay faulty data packets to other switches 802, 804, 806, 808, and 810 in the network, but will also perform data integrity checks on faulty data packets 824, 826, 828, and 830 and construct current health metrics for the corresponding ports 832 and 834 based on identified faults using CRC error counters, FEC statistics, and histogram offsets over time. When the current health metrics of ports 832 and 834 drop below the unhealthy threshold, the first switch 800 switches to SAF mode to prevent further propagation of faulty packets. In the example shown, although many faulty data packets 824, 826, 828, and 830 are received by ports 832 and 834 of the first switch 800, a more limited number of faulty data packets 824, 826, 828, and 830 are propagated to receiving ports 836 and 838 of other switches 802, 804, 806, 808, and 810 in the network.

[0052] In at least one embodiment, if a first switch 800 operating in CT mode identifies faulty data packets 824, 826, 828, and 830 after they have been propagated to other switches 802, 804, 806, 808, and 810, the first switch 800 may send one or more subsequent data packets indicating that the first switch 800 knows that the faulty data packets 824, 826, 828, and 830 are faulty. Such subsequent data packets can simplify network error diagnosis and inform those other switches 802, 804, 806, 808, and 810 of their decision on whether to switch to SAF mode for those ports 836 and 838. For example, downstream switches 802, 804, 806, 808, and 810 can also perform the same data integrity checks as the first switch 800 and receive subsequent data packets indicating that the first switch 800 knows faulty data packets 824, 826, 828, and 830. By comparing the number of self-identified faulty data packets 824, 826, 828, and 830 with the number of known faulty data packets 824, 826, 828, and 830 indicated by subsequent data packets, downstream switches 802, 804, 806, 808, and 810 can determine that the link between downstream switches 802, 804, 806, 808, and 810 and the first switch 800 (including any intermediate switches 802, 804, 806, 808, and 810) is likely healthy. In contrast, if the numbers do not match, then a faulty link or upstream switches 800, 802, 804, 806, 808, and 810 can be identified.

[0053] refer to Figure 9 This diagram illustrates the propagation of data packets through a network of switches 900, 902, 904, 906, 908, and 910 according to an exemplary embodiment of the present disclosure. With switches 900, 902, 904, 906, 908, and 910 configured for either CT mode or SAF mode, the first switch 900 can otherwise receive health data packets 924, 926, and 928 from a first port 932 with a faulty link. Based on the current health indicators of port 932, in addition to switching to SAF mode, the first switch 900 can also identify the faulty link and communicate the error to a management plane monitor. Once the faulty link is corrected, the management plane monitor can communicate the correction to the first switch 900, which can then switch back to CT mode or continue monitoring the faulty data packets.

[0054] refer to Figure 10This diagram illustrates the propagation of data packets through a network of switches 1000, 1002, 1004, 1006, 1008, and 1010 according to an exemplary embodiment of the present disclosure. With switches 1000, 1002, 1004, 1006, 1008, and 1010 configured for CT mode or SAF mode, the first switch 1000 can receive healthy data packets 1024 and 1028 and faulty data packets 1026 from the first port 1032. Based on the current health indicators of port 1032, the first switch 1000 can remain in CT mode, allowing a small number of faulty data packets 1026 to propagate to downstream switches 1002 and 1006. A small background CRC error rate does not trigger SAF mode or alert the management plane monitor in any of the 1000, 1002, 1004, 1006, 1008, or 1010 switches because the small background CRC error rate is less troublesome for the entire network compared to the latency of operating all 1000, 1002, 1004, 1006, 1008, or 1010 switches in SAF mode.

[0055] Embodiments of this disclosure can identify when a specific source or upstream device experiences a higher-than-normal number of problems. If a switch detects an error rate exceeding an expected baseline, the switch can switch to a more fault-tolerant mode and report the problem to a management plane monitor. The statistical data and recorded health metrics recorded by embodiments of this disclosure are useful for data centers where the number of human operators is relatively small compared to the number of machines and switches.

[0056] It is believed that the inventive concept and its many accompanying advantages will be understood from the foregoing description of embodiments of the inventive concept disclosed herein, and it will be appreciated that various changes can be made in the form, construction, and arrangement of its components without departing from the broad scope of the inventive concept disclosed herein or without sacrificing all its significant advantages; and individual features from various embodiments can be combined to achieve other embodiments. The forms previously described herein are merely illustrative embodiments, and the following claims are intended to cover and include these changes. Furthermore, any feature disclosed with respect to any individual embodiment may be incorporated into any other embodiment.

Claims

1. A computer device comprising: At least one processor communicates with a memory storing processor-executable code for configuring the at least one processor to: Data packets are received and relayed via multiple data ports; In pass-through CT mode, each data packet is relayed from the receive port to the relay port before the entire data packet is received. Perform one or more data integrity checks on each data group; Associate each data integrity check with the receiving port; When the port health metric associated with a specific port, as measured by a data integrity check associated with that specific port, drops below a defined unhealthy threshold, the system switches to Store & Forward (SAF) mode. When in the CT mode, identify faulty data groups; and The faulty data packet is reported to the downstream device receiving the faulty data packet using subsequent data packets, wherein the subsequent data packets indicate the faulty data packet so that the downstream device determines a faulty link based on the subsequent data packets and one or more self-identified faulty data packets, and wherein the faulty link is determined by determining a mismatch between the number of self-identified faulty data packets and the number of faulty data packets indicated by the subsequent data packets.

2. The computer device of claim 1, wherein the at least one processor is further configured to: While in the SAF mode, continue to receive and relay data packets via the plurality of data ports; and After the predetermined time period, the mode will be changed back to the CT mode.

3. The computer device of claim 1, wherein the at least one processor is further configured to: While in the SAF mode, data packets continue to be received and relayed via the plurality of data ports, and one or more data integrity checks are performed on each data packet; and When the port health metric associated with the specific port rises above the defined health threshold, the system reverts to the CT mode.

4. The computer device according to claim 1, wherein: The at least one processor is further configured to maintain a histogram of data integrity check failures for one or more of the plurality of data ports; and The port health metric includes the offset of the histogram over time.

5. The computer device of claim 1, wherein the at least one processor is further configured to operate a first port of the plurality of data ports in CT mode, while simultaneously operating a second port of the plurality of data ports in SAF mode.

6. The computer device of claim 1, wherein the at least one processor is further configured to report the specific port to a management plane monitor.

7. A network device comprising: At least one processor communicates with a memory storing processor-executable code for configuring the at least one processor to: Data packets are received and relayed via multiple data ports; In pass-through CT mode, each data packet is relayed from the receive port to the relay port before the entire data packet is received. Perform one or more data integrity checks on each data group; Associate each data integrity check with the receiving port; When the port health metric associated with a specific port, as measured by a data integrity check associated with that specific port, drops below a defined unhealthy threshold, the system switches to Store & Forward (SAF) mode. When in the CT mode, identify faulty data groups; and The faulty data packet is reported to the downstream device receiving the faulty data packet using subsequent data packets, wherein the subsequent data packets indicate the faulty data packet so that the downstream device determines a faulty link based on the subsequent data packets and one or more self-identified faulty data packets, and wherein the faulty link is determined by determining a mismatch between the number of self-identified faulty data packets and the number of faulty data packets indicated by the subsequent data packets.

8. The network apparatus of claim 7, wherein the at least one processor is further configured to: While in the SAF mode, continue to receive and relay data packets via the plurality of data ports; and After the predetermined time period, the mode will be changed back to the CT mode.

9. The network apparatus of claim 7, wherein the at least one processor is further configured to: While in the SAF mode, data packets continue to be received and relayed via the plurality of data ports, and one or more data integrity checks are performed on each data packet; and When the port health metric associated with the specific port rises above the defined health threshold, the system reverts to the CT mode.

10. The network device according to claim 7, wherein: The at least one processor is further configured to maintain a histogram of data integrity check failures for one or more of the plurality of data ports; and The port health metric includes the offset of the histogram over time.

11. The network apparatus of claim 7, wherein the at least one processor is further configured to operate a first port of the plurality of data ports in CT mode, while simultaneously operating a second port of the plurality of data ports in SAF mode.

12. The network apparatus of claim 7, wherein the at least one processor is further configured to report the specific port to a management plane monitor.

13. A network system comprising: At least one network device, comprising: At least one processor communicates with a memory storing processor-executable code for configuring the at least one processor to: Data packets are received and relayed via multiple data ports; In pass-through CT mode, each data packet is relayed from the receive port to the relay port before the entire data packet is received. Perform one or more data integrity checks on each data group; Associate each data integrity check with the receiving port; When the port health metric associated with a specific port, as measured by a data integrity check associated with that specific port, drops below a defined unhealthy threshold, the system switches to Store & Forward (SAF) mode. When in the CT mode, identify faulty data groups; and The faulty data packet is reported to the downstream device receiving the faulty data packet using subsequent data packets, wherein the subsequent data packets indicate the faulty data packet so that the downstream device determines a faulty link based on the subsequent data packets and one or more self-identified faulty data packets, and wherein the faulty link is determined by determining a mismatch between the number of self-identified faulty data packets and the number of faulty data packets indicated by the subsequent data packets.

14. The network system of claim 13, wherein the at least one processor is further configured to: While in the SAF mode, continue to receive and relay data packets via the plurality of data ports; and After the predetermined time period, the mode will be changed back to the CT mode.

15. The network system of claim 13, wherein the at least one processor is further configured to: While in the SAF mode, data packets continue to be received and relayed via the plurality of data ports, and one or more data integrity checks are performed on each data packet; and When the port health metric associated with the specific port rises above the defined health threshold, the system reverts to the CT mode.

16. The network system according to claim 13, wherein: The at least one processor is further configured to maintain a histogram of data integrity check failures for one or more of the plurality of data ports; and The port health metric includes the offset of the histogram over time.

17. The network system of claim 13, wherein the at least one processor is further configured to operate a first port of the plurality of data ports in CT mode, while simultaneously operating a second port of the plurality of data ports in SAF mode.

Citation Information

Patent Citations

  • Network method, network apparatus, and computer readable storage medium

    CN108234306A

  • Normalization of detecting and reporting failures for a memory device

    CN111383707A