Mechanism to identify link breakage cause

By setting parameters on the LFA and PTM switching ports, the cause of network link interruption can be identified, solving the problem of cumbersome link fault debugging in existing technologies and achieving fast and accurate fault troubleshooting.

CN116471165BActive Publication Date: 2026-08-04NVIDIA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NVIDIA CORP
Filing Date
2022-12-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In network deployment, there is a lack of simple and effective mechanisms for identifying the causes of link interruptions, and the debugging process is cumbersome and requires professional knowledge.

Method used

The Link Failure Analyzer (LFA) uses the Topology Manager (PTM) to exchange port setting parameters, identifies port setting mismatches and sends alarm messages, and combines this with the topology file to determine the cause of the link failure.

Benefits of technology

Quickly identify the cause of link interruption, reduce debugging time, avoid reliance on professional knowledge, and improve troubleshooting efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116471165B_ABST
    Figure CN116471165B_ABST
Patent Text Reader

Abstract

The present disclosure relates to mechanisms for identifying link outage causes. Methods, systems, and devices for mechanisms for identifying link outage causes are provided herein. As described herein, a first port of a first peer device can be determined to have unexpectedly changed to a port outage state. Subsequently, a topology file can be referenced to identify a second port of a second peer device with which the first peer device would have a link if not for the first port being in the port outage state. In some examples, port settings of the first port can be compared to port settings of the second port. If the port settings of the first port do not match the associated port settings of the second port, an alert message can be sent to a network administrator indicating that such a mismatch is a likely cause of the first port being in the port outage state.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to systems, methods, and apparatus for troubleshooting connectivity errors between network devices, and more specifically to mechanisms for identifying the causes of link interruptions. Background Technology

[0002] In network deployments, there is no simple mechanism to identify the cause of a particular link failure. A link may involve two (2) ports connected at both ends. Each port is programmed with various settings, such as speed, auto-negotiation, forward error correction (FEC), interface type, etc., and the settings at both ends of the link must be compatible (e.g., the port settings at the first end of the link must be compatible with the port settings at the second end of the link). Even a small error in the programming settings, or a small incompatibility between the port settings at either end of the link, can cause a link failure or failure (e.g., the link is not connected). Debugging and finding the root cause of a link failure (e.g., the cause of a link failure) is a tedious task and may require domain expertise. Summary of the Invention

[0003] Examples of aspects of this disclosure include:

[0004] A method includes: determining that a first port of a first peer device has unexpectedly become a port interrupted state; referring to a topology file to identify a second port of a second peer device, if the first peer device intends to have a link with the second peer device if it were not because the first port was in a port interrupted state; determining that the port settings of the first port do not match the associated port settings of the second port; and in response to determining that the port settings of the first port do not match the associated port settings of the second port, sending an alert message to a network administrator.

[0005] In any aspect of this document, the alarm message includes a description of possible causes for the port interruption status.

[0006] In any aspect of this document, the description of possible causes of the port interruption state includes a description of a port setting mismatch between the first port and the second port.

[0007] In any aspect of this document, the first peer device includes a first link failure analyzer (LFA), the second peer device includes a second LFA, and the first LFA communicates with the second LFA via a management network to exchange port settings.

[0008] In any aspect of this document, the topology file is maintained by a specified topology manager (PTM), the first LFA references the topology file through the PTM, and the second LFA also references the topology file through the PTM.

[0009] In any aspect of this document, a Link Failure Analyzer (LFA) is provided at one of the first peer device and the second peer device, wherein the topology file is maintained by the PTM, and wherein the LFA references the topology file through the PTM.

[0010] In any aspect of this document, the LFA is configured to obtain hostname and interface name information from the topology file.

[0011] Any aspect of this document further includes: determining at the LFA that the port settings of the first port match the port settings of the second port; and including information describing cable or connection problems as possible sources of link failure.

[0012] In any aspect of this document, the LFA includes a first LFA provided at the first peer device, and a second LFA provided at the second peer device, and further includes: enabling the first LFA and the second LFA to exchange port settings of the first port and port settings of the second port with each other; and configuring the first LFA to compare the port settings of the first port with the port settings of the second port to determine that the port settings of the first port match the port settings of the second port.

[0013] In any aspect of this document, the alarm message includes a description of possible causes of the link failure.

[0014] A system includes: a processor; and a memory coupled to and readable by the processor, wherein instructions are stored therein, which, when executed by the processor, cause the processor to: determine that a first port of a first peer device has been unexpectedly changed to a port interruption state; refer to a topology file to identify a second port of a second peer device, and if the first peer device intends to have a link with the second peer device, excluding the fact that the first port is in a port interruption state; determine that the port settings of the first port do not match the associated port settings of the second port; and in response to determining that the port settings of the first port do not match the associated port settings of the second port, send an alert message to a network administrator.

[0015] In any aspect of this document, the alarm message includes a description of possible causes for the port interruption status.

[0016] In any aspect of this document, the description of possible causes of the port interruption state includes a description of a port setting mismatch between the first port and the second port.

[0017] In any aspect of this document, the first peer device includes a first LFA, the second peer device includes a second LFA, and the first LFA communicates with the second LFA via a management network to exchange port settings.

[0018] In any aspect of this document, the topology file is maintained by a PTM, the first LFA references the topology file through the PTM, and the second LFA also references the topology file through the PTM.

[0019] In any aspect of this document, an LFA is provided at one of the first peer device and the second peer device, wherein the topology file is maintained by a PTM, and wherein the LFA references the topology file through the PTM.

[0020] In any aspect of this document, the LFA is configured to obtain hostname and interface name information from the topology file.

[0021] In any aspect of this document, the instructions further cause the processor to: determine at the LFA that the port settings of the first port match the port settings of the second port; and include information describing cable or connection problems as possible sources of link failure.

[0022] In any aspect of this document, wherein the LFA includes a first LFA provided at the first peer device, and wherein a second LFA is provided at the second peer device, the instructions further cause the processor to: enable the first LFA and the second LFA to exchange port settings of the first port and port settings of the second port with each other; and configure the first LFA to compare the port settings of the first port with the port settings of the second port to determine that the port settings of the first port match the port settings of the second port.

[0023] A first peer device includes: a plurality of ports; an application program; a processor; and a memory coupled to and readable by the processor, wherein instructions are stored, when executed by the processor, causing the processor to: determine that a first port among the plurality of ports of the first peer device has unexpectedly become a port interrupted state; identify a second port of a second peer device by referring to a topology file via the application program, and if the first peer device intends to have a link with the second peer device, provided that the first peer device is not in a port interrupted state; determine that the port settings of the first port do not match the associated port settings of the second port; and, in response to determining that the port settings of the first port do not match the associated port settings of the second port, send an alert message to a network administrator.

[0024] Any aspect in combination with one or more other aspects.

[0025] Any one or more features disclosed herein.

[0026] Basically any one or more features disclosed herein.

[0027] Combination of any one or more features substantially disclosed herein with any one or more other features substantially disclosed herein.

[0028] Any aspect / feature / embodiment of the present invention in combination with any or more other aspects / features / embodiments.

[0029] Use of any or more aspects or features disclosed herein.

[0030] It should be understood that any feature described herein may be claimed in combination with any other feature described herein, regardless of whether such features are derived from embodiments of the same description.

[0031] Details of one or more aspects of this disclosure are set forth in the accompanying drawings and the description below. Other features, objects, and advantages of the technology described in this disclosure will be apparent from the description, the drawings, and the claims.

[0032] The phrases “at least one,” “one or more,” and “and / or” are open-ended expressions that are both connective and disconnective in operation. For example, each of the expressions “at least one of A, B, and C,” “at least one of A, B, or C,” “one or more of A, B, and C,” “one or more of A, B, or C,” and “A, B, and / or C” refers to A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together. When each of A, B, and C in the above expressions refers to an element, such as X, Y, and Z, or a class of elements, such as X1-Xn, Y1-Ym, and Z1-Zo, the phrase is intended to refer to a single element selected from X, Y, and Z, a combination of elements selected from the same class (such as X1 and X2), or a combination of elements selected from two or more classes (such as Y1 and Zo).

[0033] The term "a" or "an" refers to one or more of the same entity. Therefore, the terms "a" (or "one"), "one or more," and "at least one" are used interchangeably herein. It should also be noted that the terms "comprising," "including," and "having" are used interchangeably.

[0034] The foregoing is a simplified summary of this disclosure to provide an understanding of certain aspects of it. This summary is neither a broad nor exhaustive overview of this disclosure, its aspects, embodiments, and configurations. It is not intended to identify key or essential elements of this disclosure, nor to define its scope, but rather to present selected concepts of this disclosure in a simplified form as an introduction to the more detailed description that follows. As will be understood, other aspects, embodiments, and configurations of this disclosure may utilize one or more of the features described above, either alone or in combination, or in detail below.

[0035] Many additional features and advantages are described herein, and will be apparent to those skilled in the art when taken into account the following detailed description and the figures. Attached Figure Description

[0036] The accompanying drawings, incorporated in and forming part of this specification, illustrate several examples of the contents of this disclosure. These drawings, together with the description, explain the principles of this disclosure. The drawings simply illustrate preferred and alternative examples of how this disclosure can be made and used, and should not be construed as limiting this disclosure solely to the illustrated and described examples. Further features and advantages will become apparent from the following more detailed description of various aspects, embodiments, and configurations of this disclosure, as illustrated by the accompanying drawings mentioned below.

[0037] This disclosure is described in conjunction with the accompanying figures, which are not necessarily drawn to scale.

[0038] Figure 1 A block diagram illustrating a network system according to at least one exemplary embodiment of the present disclosure is provided.

[0039] Figure 2 A first flowchart illustrating an embodiment of the present disclosure is provided.

[0040] Figure 3 A second flowchart illustrating an embodiment of the present disclosure is provided.

[0041] Figure 4 A third flowchart illustrating an embodiment of this disclosure is provided; and

[0042] Figure 5 A fourth flowchart illustrating an embodiment of the present disclosure is provided. Detailed Implementation

[0043] It should be understood that the aspects disclosed herein can be combined in ways other than those specifically presented in the description and figures. It should also be understood that, depending on the examples or embodiments, certain actions or events of any process or method described herein may be performed in a different order, and / or may be added, combined, or not performed at all (e.g., according to different embodiments of this disclosure, all described actions or events may not be necessary to perform the disclosed technology). Furthermore, although some aspects of this disclosure are described as being performed by a single module or unit for clarity, it should be understood that the technology of this disclosure can be performed by a combination of units or modules associated with, for example, computing devices and / or medical devices.

[0044] In one or more examples, the methods, processes, and techniques may be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the function may be stored as one or more instructions or code on a computer-readable medium and executed by a hardware-based processing unit. Additionally, the function may be implemented using machine learning models, neural networks, artificial neural networks, or combinations thereof (alone or in combination with instructions). Computer-readable media may include non-transitory computer-readable media, which correspond to tangible media such as data storage media (e.g., random access memory (RAM), read-only memory (ROM), electrically erasable programmable ROM (EEPROM), flash memory, or any other medium that can be used to store desired program code in the form of instructions or data structures and is accessible by a computer).

[0045] Instructions may be executed by one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors (e.g., Intel Core i3, i5, i7, or i9 processors; Intel Celeron processors; Intel Xeon processors; Intel Pentium processors; AMD Ryzen processors; AMD Athlon processors; AMD Phenom processors; Apple A10 or 10X Fusion processors; Apple A11, A12, A12X, A12Z, or A13 Bionic processors; or any other general-purpose microprocessors), graphics processing units (e.g., Nvidia GeForce RTX 2000 series processors, Nvidia GeForce RTX 3000 series processors, AMD Radeon RX 5000 series processors, AMD Radeon RX 6000 series processors, or any other graphics processing units), application-specific integrated circuits (ASICs), field-programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Therefore, the term "processor" as used herein may refer to any of the above-described structures or any other physical structure suitable for implementing the described techniques. Furthermore, these techniques may be fully implemented in one or more circuit or logic elements.

[0046] Before explaining any embodiment of this disclosure in detail, it should be understood that this disclosure, in its application, is not limited to the structural details and arrangement of components listed in the following description or illustrated in the accompanying drawings. This disclosure can be implemented in other ways and can be carried out or performed in various manner. Furthermore, it should be understood that the phrases and terms used herein are for descriptive purposes only and should not be considered limiting. The terms “comprising,” “including,” or “having,” and variations thereof, as used herein, refer to items listed herein and their equivalents, as well as other items. In addition, examples may be used to illustrate one or more aspects of this disclosure. Unless otherwise expressly stated, the use or enumeration of one or more examples (which may be expressed in terms of “for example,” “example,” “e.g.,” or similar language) is not intended to, nor to, limit the scope of this disclosure.

[0047] The following description provides only embodiments and is not intended to limit the scope, applicability, or configuration of the claims. Rather, it will provide those skilled in the art with an advantageous description for implementing the embodiments. It will be understood that various changes may be made to the function and arrangement of the elements without departing from the spirit and scope of the appended claims. Unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains. It will be further understood that terms, such as those defined in common dictionaries, should be interpreted as having the same meaning as they have in the relevant art and in this disclosure.

[0048] As can be understood from the following description, and for computational efficiency reasons, system components can be placed in any suitable location in a distributed network of components without affecting the operation of the system.

[0049] Furthermore, it should be understood that the various links connecting the components can be wired, traced, or wireless links, or any suitable combination thereof, or any other suitable known or later-developed element capable of providing data to and / or communicating with the connected components. For example, the transmission medium used as a link can be any suitable carrier of electrical signals, including coaxial cables, copper wires and optical fibers, electrical traces on printed circuit boards (PCBs), or the like.

[0050] The terms “determine,” “calculate,” and “compute,” as used herein, and their variations, are used interchangeably and include any appropriate type of method, process, operation, or technique.

[0051] This document will describe various aspects of this disclosure with reference to the accompanying drawings, which can be used as schematic diagrams of an idealized configuration.

[0052] In network deployments, there is no simple mechanism to identify the cause of a particular link failure. A link may involve two (2) ports connected at both ends. Each port is programmed with multiple settings, such as speed, auto-negotiation, forward error correction (FEC), interface type, etc., and the settings at both ends of the link must be compatible (e.g., the port settings at the first end of the link must be compatible with the port settings at the second end of the link). Even a small error in the programming settings, or a small incompatibility between the port settings at both ends of the link, can cause a link failure or interruption (e.g., the link is not connected). Debugging and finding the root cause of a link failure (e.g., the cause of a link interruption) is a tedious task and may require domain expertise.

[0053] The problem with link failures is that, unlike other failure scenarios (such as protocol session interruptions), the cause of a failure can usually be narrowed down because there is data exchange between network devices through the linked ports (e.g., the network devices and / or ports to be linked can be referenced as peers or peer devices). However, in the case of a link failure, since the link itself has not been established, it is impossible to narrow down the cause of the link interruption due to a lack of communication. Some low-level hardware mechanisms can be used to vaguely indicate what the problem or source of the link failure might be, but these low-level hardware mechanisms may be difficult to detect (e.g., the mechanism is not exposed) and may not be able to accurately or specifically identify the root cause of the problem (e.g., mismatched configurations).

[0054] The embodiments proposed herein were conceived with the above and other issues in mind.

[0055] The inventive concept relates to mechanisms for identifying the causes of link interruptions. For example, an application on a first network device (e.g., a Link Failure Analyzer (LFA) application) as described herein can leverage the use of a topology manager (e.g., a Prescribed Topology Manager (PTM)) to exchange link-based parameters (e.g., with a corresponding application on a peer device intended to link to the first network device) to determine the cause of the link interruption (e.g., the cause of a link interruption or failure). In some cases, the PTM might be an inbox solution that uses a topology file to identify the two ends of a proposed link, so that when the link appears (e.g., the two connections of the link have been plugged into the respective ports of the network devices intended to be linked), the PTM can verify the link (e.g., using a Link Layer Discovery Protocol (LLDP)) to see if the connection was correctly established (e.g., based on information included in the topology file). However, this verification may only work if the link is connected or has been active at all times.

[0056] In some examples, the topology manager can access additional information about the respective sides or ends of the proposed link (e.g., through a topology file maintained by the topology manager). For instance, the topology manager can obtain information about one or more ports of a first network device on a first side or end (e.g., source / side) of the proposed link, and one or more ports of a second network device on a second side or end (e.g., remote or peer / side) of the same proposed link. In some examples, the topology manager can determine this information from a topology file maintained by the topology manager, where the topology file indicates or lists which ports of which network devices are proposed to have links, along with additional information about the ports and network devices to be linked (e.g., hostname, interface name, etc. for each port or network device).

[0057] Based on this information in the topology file and maintained by the topology manager, user applications (e.g., LFA applications) can be generated (e.g., one or two network devices at either end of a proposed link). The user application can resolve the initialized and configured port list to identify which ports should be up but are down. Ports considered up but down can be obtained from the managed ports configured to be up. After obtaining the list of unexpectedly down ports, the user application checks the topology file in the topology manager to find information about peer ports and devices with proposed links to these unexpected ports. This information about peer ports and devices obtained from the topology file may include hostnames, interface names, and other information. Based on this information (e.g., hostnames), the user application can establish a Transmission Control Protocol (TCP) session to communicate (e.g., through its management network) with the corresponding user application on the network device (e.g., a remote device) to which the unexpectedly down port is to be connected. Setting parameters and configuration information (e.g., setting DNS, modifying Access Control List (ACL) rules, etc.) to enable TCP sessions and communication between network devices may be beyond the scope of this disclosure.

[0058] Once a TCP session is established between user applications at either end of the proposed link (e.g., on two peer devices), the user applications can exchange link configuration parameters to configure the port of the proposed link. Link configuration parameters may include administrator status, speed, number of channels / breakthrough mode, auto-negotiation, forward error correction (FEC), interface type, port error status (e.g., if applicable), etc., through which the port of the proposed link has been configured. This list of link configuration parameters is not an exhaustive list of possible parameters that a port can be configured with; additional link parameters can be configured for a port. For example, some link parameters may be vendor-specific to the network equipment manufacturer. Some listed parameters (e.g., FEC, auto-negotiation, speed, etc.) may have administrative values ​​but different operational values ​​depending on the specific device firmware implementation. In this case, the user applications can exchange both administrative and operational values.

[0059] After exchanging link configuration parameters for the ports used to configure the proposed link (e.g., after obtaining information from the peer device), the user application of the network device with the unexpectedly interrupted port cross-checks the exchanged link configuration parameters to determine whether the same parameters configured on the ports at both ends of the proposed link are compatible (e.g., on the remote port and the local interrupted port).

[0060] If the parameters are compatible, the user application may take no action. Alternatively, if some parameters are incompatible, the user application can log these incompatible parameters and issue an alert. Alerts can be issued via Simple Network Management Protocol (SNMP), telemetry, or any other available alerting method. Furthermore, a command-line interface (CLI) can be provided to display commands to highlight mismatched configurations. For example, an alert message can be sent to the network administrator, which may include a description of possible causes for port outages (e.g., a description of a port configuration mismatch between the ports at both ends of a proposed link).

[0061] The example below shows a CLI display command (e.g., an alert message) output by a user application to indicate a configuration mismatch.

[0062] CLI command output example showing configuration mismatch

[0063] Ethernet 1 / 1 ToR2-Ethernet 1 / 1 FEC RS none

[0064] This example CLI display command indicates a configuration mismatch between a FEC parameter configured with a first value (e.g., Reed Solomon(RS)FEC) on a local port (e.g., the local port of a first network device) and the same FEC parameter configured with a second value (e.g., "None") on a remote port (e.g., the remote port of a second network device). Therefore, the user can receive this CLI display command output and quickly identify that the link interruption between the local and remote ports may be due to a mismatch in the FEC parameters configured for their respective ports, which the user can correct (e.g., by setting the FEC parameters at both ends of the link to the same or compatible values).

[0065] Because the mechanism for identifying the cause of link interruption described herein is an application-level solution, it can be deployed in any network device, such as servers, storage devices, intelligent network interface controllers (NICs), and products from other vendors. In some instances, the mechanism may require the user application (e.g., an LFA application) to have a plugin for obtaining port-level settings (e.g., transport configuration parameters for the port used to determine the link).

[0066] Therefore, using a topology manager (e.g., PTM), the user application of the first network device can effectively communicate with a second network device that intends to have a link with the first network device (e.g., a peer device), and can obtain the required configuration and settings of the ports of the second network device for cross-comparison with the configuration and settings of the ports of the first network device. By cross-comparing the configuration and settings of the ports at both ends of the proposed link, the user application can identify or narrow down the cause of the link failure (e.g., the cause of the link break or the reason for the link interruption). However, this mechanism assumes that the link is correctly connected and misconfigured, but there are other situations that may cause link failure (e.g., leading to a link interruption).

[0067] In some examples, another scenario leading to link failure (or interruption) might include connection errors and link disconnection. For instance, if the configuration of a second network device (e.g., a peer device) matches or is compatible with the configuration of the first network device (e.g., as listed in a topology file maintained by a topology manager / PTM), a user application (e.g., an LFA application) can identify and display the link failure as a "connection / cable problem," indicating that the configurations are compatible, but a connection or cable issue caused the link failure or interruption. Alternatively, if there is a configuration mismatch between the configuration and settings at either end of the proposed link, the user application can highlight the mismatch (e.g., send an alert message describing a port setting mismatch between ports), and if the user corrects the mismatch but the link is still not connected, the user application can identify and display the link failure as a "connection / cable problem."

[0068] In some examples, another scenario leading to a link failure (or interruption) might include situations where the connection and configuration are correct and compatible, but the link remains unconnected. In this case, the user application can identify and display the link failure as "Connection / Cable Problem". Additionally, or alternatively, another scenario leading to a link failure (or interruption) might include the second network device (e.g., a peer device) being unreachable. For example, the second network device might be unreachable due to a software failure, hardware failure, or a restart. In this case, the user application can identify and display the link interruption as "Peer Device Unreachable". Furthermore, or alternatively, another scenario leading to a link failure (or interruption) might include situations where one end of the proposed link is connected, while the other end is not connected or is loosely connected. In this case, the user application can identify and display the link failure as "Peer Cable Not Connected" or "Peer Unavailable".

[0069] Although a trigger (e.g., for the user application to reach a peer) at the first network device used to establish communication (e.g., a TCP session) with the second network device is identifying an unexpectedly interrupted port, the user application can respond to any query on that specific port, regardless of its state (e.g., connected or disconnected). By allowing the user application to respond to any query on a specific port, the user application can assist in identifying other causes of link failure or interruption, such as transceiver problems and loose connections.

[0070] Embodiments of this disclosure provide technical solutions to one or more of the following problems: (1) determining the cause of a link interruption; (2) prolonged debugging time for determining the cause of a link interruption; and (3) the need for domain expertise when attempting to determine the cause of a link interruption. For example, the user applications and mechanisms described herein (e.g., LFA applications and / or mechanisms) can be used to narrow down the causes of link failures or link interruptions without requiring the link to be connected (e.g., this was previously impossible because the other end of the link (the remote end) may be unknown before the link is connected or linked). Furthermore, the mechanisms and techniques described herein can leverage a topology manager (e.g., PTM) to establish out-of-band connections with peer devices to exchange link configuration parameters and narrow down the root cause of the link interruption, displaying the identified cause to the user. In some examples, such mechanisms for identifying the cause of a link interruption can reduce debugging time and can avoid the need for expert intervention to find the cause of the link interruption (e.g., the cause of the failure) based on narrowing down the causes of the link interruption.

[0071] First turn Figure 1 The diagram illustrates a block diagram of a system 100 according to at least one embodiment of the present disclosure. The system 100 can be used to identify the cause of a link interruption between ports of network devices that have been determined to be unexpectedly interrupted. For example, a topology file may include information about which ports of which devices are to be linked, and specific information about these ports (e.g., hostname, interface name, etc.). This specific information can then be used to establish a connection between corresponding applications on the network devices at both ends of the unexpectedly interrupted proposed link. These applications can then exchange link configuration parameters, which specify the ports configured at their respective network devices. After exchanging parameters, the applications can cross-check the configuration values ​​of each parameter for the ports at both ends of the proposed link to determine if there are configuration mismatches that could be the cause of the unexpected link interruption. In either case (e.g., a configuration mismatch exists or these are not configuration mismatches), the applications can send an alert message to the network administrator indicating possible causes of the link interruption, such as a configuration mismatch (e.g., if applicable), a "connection / cable problem" cause, a "peer unreachable" cause, a "peer cable not connected" cause, etc.

[0072] System 100 includes network device 104, communication network 108, and network device 112. In at least one exemplary embodiment, network devices 104 and 112 may correspond to network switches (e.g., Ethernet switches), a collection of network switches, a NIC, or any other suitable device for controlling data flow between devices connected to communication network 108. Each network device 104 and 112 may be connected to one or more personal computers (PCs), laptops, tablets, smartphones, servers, collections of servers, etc. In a specific but non-limiting example, each network device 104 and 112 includes multiple network switches in a fixed or modular configuration.

[0073] Examples of communication networks 108 that can be used to connect network devices 104 and 112 include Internet Protocol (IP) networks, Ethernet networks, InfiniBand (IB) networks, Fibre Channel networks, the Internet, cellular communication networks, wireless communication networks, combinations thereof (e.g., Fibre Channel over Ethernet), variations thereof, and / or the like. In a specific but non-limiting example, communication network 108 is a network that enables communication between network devices 104 and 112 using Ethernet technology. In one specific but non-limiting example, network devices 104 and 112 correspond to peer devices as described in more detail below.

[0074] Although not explicitly shown, network device 104 and / or network device 112 may include storage devices and / or processing circuitry for performing computational tasks, such as those associated with controlling data flow within each network device 104 and 112 and / or on the communication network 108. Such processing circuitry may include software, hardware, or a combination thereof. For example, processing circuitry may include memory (which contains executable instructions) and a processor (e.g., a microprocessor) that executes the instructions on the memory. Memory may correspond to any suitable type of memory device or collection of memory devices configured to store instructions. Non-limiting examples of suitable memory devices that may be used include flash memory, random access memory (RAM), read-only memory (ROM), variations thereof, combinations thereof, or the like.

[0075] In some embodiments, network device 104 and / or network device 112 may include an application stored in memory, such as LFA application 116 or LFA application 120, which is induced and responsible for supporting the mechanism for identifying the cause of the link outage described herein. As illustrated, in some examples, LFA applications 116, 120 may be induced or may otherwise be in or accessed by network devices 104, 112 (e.g., LFA applications 116, 120 are loaded or programmed onto network devices 104, 112). Additionally or alternatively, LFA applications 116, 120 may be external to network devices 104, 112 (i.e., included within network devices 104, 112). Figure 1 (In some other hardware components of System 100).

[0076] In some embodiments, the LFA application 116, 120 at any network device 104, 112 may be launched after generating a list of initialized and configured ports and determining that one or more ports of network devices 104, 112 in the list are expected to be connected but are interrupted (e.g., link interrupted). Alternatively, the LFA application 116, 120 may be launched or induced in response to any query for a particular port, regardless of the port's state (e.g., connected or disconnected). The LFA application 116, 120 at network devices 104, 112 may then identify the corresponding network device 104, 112 and / or the corresponding port that should be connected to the link interrupted port based on a reference to topology file 128. Topology file 128 may be maintained by a topology manager 124 (e.g., PTM), which is provided within one of network devices 104, 112 and / or within a switch in the communication network 108.

[0077] As a non-limiting example, the LFA application 116 provided at network device 104 can be configured to identify that network device 104 should be connected to network device 112 based on information in topology file 128 (e.g., via topology manager 124 reference). Based on the assumption that network devices 104 and 112 should have a connection or link, network devices 104 and 112 can be referred to as peer devices. In some examples, the LFA application 116 can communicate directly with the LFA application 120 provided at network device 112 to exchange information (e.g., link configuration parameters) to ensure that the ports at the network devices 104 and 112 to be connected (e.g., based on information from topology file 128) are configured with compatible parameters. Alternatively, the LFA application 116 may determine information and / or configuration parameters (e.g., link configuration parameters of the ports of network device 112) of network device 112 through other components or devices of system 100 (e.g., topology manager 124, management device, management network, etc.) to compare the configuration parameters between network device 104 and network device 112 (e.g., when determining the cause of a link interruption).

[0078] Because this mechanism for identifying the cause of link interruption is an application-level solution (e.g., using LFA applications 116, 120), LFA applications 116, 120 can be deployed in any device of system 100, such as servers, storage devices, smart NICs, and products from other vendors. Furthermore, the mechanism for identifying the cause of link interruption described herein may require LFA applications 116, 120 to have plugins for obtaining port-level settings (e.g., from topology file 128).

[0079] To some extent, topology file 128 can indicate or list which ports of which network devices are intended to have links. Furthermore, topology file 128 can include additional information about the ports and network devices to be linked, such as the hostname, interface name, etc., of each port or network device. Based on the information in topology file 128 (e.g., accessed via topology manager 124), LFA applications 116, 120 provided on their respective network devices can form TCP sessions to communicate with each other (e.g., through network management). Subsequently, using the TCP sessions, LFA applications 116, 120 can exchange link configuration parameters for the ports and network devices to be linked to determine if the configuration parameters of each port to be linked are compatible (e.g., as part of determining the cause of a link failure). Therefore, using topology manager 124, LFA applications 116, 120 can effectively communicate and exchange information to obtain the configuration and settings of the ports to be linked, cross-compare the configurations and settings, and identify the cause of the link failure.

[0080] The topology file 128 can be maintained at a centralized PTM, or it can be partially maintained at several PTMs, which can be distributed across one or more network devices 104, 112. For example, the topology manager 124 can be provided at a centralized controller in system 100, or it can be distributed across several network devices 104, 112.

[0081] In addition to maintaining the topology file 128, the topology manager 124 can provide other functions or operations for the system 100. For example, the topology manager 124 can act as a dynamic cabling verification tool to help detect and eliminate connection errors. For example, the topology manager 124 can adopt a specified network cabling plan (e.g., the contents of a topology.dot file generated and stored by many operators) and can combine the cabling plan with runtime information (e.g., information obtained from the Link Layer Discovery Protocol (LLDP)) to verify that the actual cabling and connections in the system 100 match the cabling plan. To verify that the actual cabling and connections match the cabling plan, the topology manager 124 can communicate with different components of the system 100 (e.g., network device 104, communication network 108, network device 112, etc.) to determine which network devices are linked, which specific ports are linked, or other links in the system 100, and can check these determined links against the cabling plan.

[0082] In some embodiments, the memory and processor of network devices 104, 112 may be integrated into a common device (e.g., a microprocessor may include integrated memory). Alternatively or additionally, the processing circuitry may include hardware such as an application-specific integrated circuit (ASIC). Other non-limiting examples of processing circuitry include integrated circuit (IC) chips, central processing units (CPUs), general-purpose processing units (GPUs), microprocessors, field-programmable gate arrays (FPGAs), collections of logic gates or transistors, resistors, capacitors, inductors, diodes, or the like. Some or all of the processing circuitry may be provided on a single PCB or an array of PCBs. It should be understood that any suitable type of electrical component or collection of electrical components may be adapted to incorporate into the processing circuitry.

[0083] Furthermore, although not explicitly shown, it should be understood that network devices 104 and 112 include one or more communication interfaces for facilitating wired and / or wireless communication between each other and other components of system 100 not shown.

[0084] Figure 2A flowchart 200 illustrates various aspects of this disclosure. Flowchart 200 may include different steps or operations for identifying mechanisms for identifying link outage causes as described herein. The steps and operations of flowchart 200 may be performed, for example, by at least one processor or otherwise. The at least one processor may be the same as or similar to the processor described above. The at least one processor may be part of a computing device (such as a personal computer, laptop, smartphone, etc.) or part of the network device or peer device described above. For example, the steps and operations of flowchart 200 may be performed by an application described herein (e.g., an LFA application) and / or a topology manager (e.g., a PTM) or otherwise.

[0085] At operation 204 of flowchart 200, the application of the first network device (e.g., an LFA application) can identify a list of unexpectedly interrupted ports of the first network device. For example, the application can obtain this list of unexpectedly interrupted ports from a topology file maintained by a topology manager (e.g., a PTM), allowing the application to reference the topology file via the topology manager. In some examples, the application can determine which ports are unexpectedly interrupted (e.g., those expected to be connected but interrupted) based on the ports that are managed to be connected. Therefore, unless a port is interrupted for a specific reason (e.g., being offline, a known problem, etc., which can be indicated in the topology file), the application can determine that a port indicated as interrupted in the topology file is unexpectedly interrupted. Furthermore, the list of unexpectedly interrupted ports can include multiple unexpectedly interrupted ports, and the remaining operations of flowchart 200 (e.g., operations 208 through 240) can be performed for each port in the list of unexpectedly interrupted ports. That is, the mechanism described herein for identifying the cause of link interruption can be used or performed for each individual port determined to be unexpectedly interrupted.

[0086] In operation 208, the application can obtain peering information for a given port that was unexpectedly interrupted from the topology file. For example, the application can access the topology file (e.g., via a topology manager or PTM) to identify a specific port of a second network device to which the given unexpectedly interrupted port is intended to have a link (e.g., if the link were not interrupted). This specific port and / or the second network device may be referred to as a peer of the first network device. In some examples, the first network device may be referred to as the first peer, and the second network device may be referred to as the second network device (e.g., based on the first and second network devices having an intended link becoming peers of each other). As part of the peering information collected by the application, the application may also determine additional information about the specific port, such as hostname, interface name, etc.

[0087] In operation 212, the application at the first network device can attempt to locate the second network device using the peering information collected at operation 208. For example, the application can use additional information about a specific port maintained in the topology file (e.g., the hostname of the second network device, etc.) to try and reach the second network device in the system.

[0088] In operation 216, if the application cannot reach the second network device, the application can identify and display the link interruption reason for the given unexpectedly interrupted port as "peer unavailable" or "peer cable not connected." That is, based on the inability to reach the second network device, the application at the first network device can identify that the port and / or the intended link for that port is interrupted because the second network device is offline or unavailable for other reasons (e.g., because the end of the intended link at the second network device is not connected or is loosely connected). After identifying and displaying the link interruption reason as part of operation 216, flowchart 200 can return to operation 208, and the application at the first network device can perform the operations of flowchart 200 starting from operation 208 to identify the link interruption reason for the next unexpectedly interrupted port in the list identified from operation 204.

[0089] Alternatively, in operation 220, if the application at the first network device can find the second network device, the application may attempt to establish a peer session with the corresponding application at the second network device (e.g., through network management). For example, the application at the first network device may attempt to establish a TCP session to enable communication with the corresponding application at the second network device. In some examples, the application at the first network device may be referred to as the first peer device, and the corresponding application at the second network device may be referred to as the second peer device.

[0090] In operation 224, if the application is unable to establish a peering session, it can identify and display the link interruption cause for the given unexpectedly interrupted port as "peer unreachable". That is, based on the ability to find the second network device but the inability to establish a peering session, the application at the first network device can identify that the port and / or the intended link for that port is interrupted because the second network device is unreachable (e.g., due to software failure, hardware failure, or a restart of the second network device). After identifying and displaying the link interruption cause as part of operation 224, flowchart 200 can return to operation 208, and the application at the first network device can perform the operations of flowchart 200 that began in operation 208 to identify the link interruption cause for the next unexpectedly interrupted port in the list identified in operation 204.

[0091] Alternatively, in operation 228, if the application at the first network device can establish a peering session with the corresponding application at the second network device, the applications can exchange peering port settings (e.g., link configuration parameters). For example, the application at the first network device can receive port settings already configured for a specific port from the corresponding application at the second network device through the peering session, with a given unexpectedly interrupted port intended to have a link with that specific port. These peering port settings may include administrator status, speed, number of channels / breakthrough mode, auto-negotiation, forward error correction (FEC), interface type, port error status (e.g., if applicable), etc., through which a specific port of the second network device has been configured. This list of peering port settings is not intended to be an exhaustive list of possible settings or parameters that can be configured for a port, and ports can be configured with additional port settings. For example, some port settings may be vendor-specific for the network device manufacturer. Furthermore, some listed parameters (e.g., FEC, auto-negotiation, speed, etc.) may have administrative values, but different operational values ​​depending on the specific device firmware implementation. In this case, the applications can exchange both administrative and operational values.

[0092] In operation 232, the application at the first network device can compare the settings configured for the unexpectedly interrupted port at the first network device with the peer port settings of the port (e.g., peer port) of the second network device that is exchanged and received at operation 228 to determine whether there are any mismatches or incompatibilities between the settings of the two ports. For example, the application can check whether the unexpectedly interrupted port at the first network device and the port at the second network device that is intended to have a link with the unexpectedly interrupted port have compatible values ​​for each configured setting (e.g., the same value or a value compatible for enabling communication).

[0093] In operation 236, if the application determines that there is no configuration mismatch between the unexpectedly interrupted port and the corresponding port used to establish the link at the second network device (e.g., all settings for each port are compatible for enabling communication between the ports), the application can identify and display the link interruption cause for that given unexpectedly interrupted port as a "cable / connection problem." In some examples, the cause of a "cable / connection problem" could indicate an incorrect connection between ports, port configurations and settings being compatible but a problem with the connection or cable causing the link failure or interruption, a configuration mismatch being identified and corrected but the link still being interrupted, a connection and configuration being correct and compatible but the link still being interrupted, or a combination thereof. After identifying and displaying the link interruption cause as part of operation 236, flowchart 200 can return to operation 208, and the application at the first network device can perform the operations of flowchart 200, which began in operation 208, to identify the link interruption cause for the next unexpectedly interrupted port in the list identified from operation 204.

[0094] Alternatively, in operation 240, if the application does determine that there is at least one configuration mismatch between the settings of the port that was unexpectedly interrupted and the settings of the corresponding port used to establish the link on the second network device, the application may log the mismatch and issue a warning to the network administrator or other user regarding the mismatched configuration. That is, if the application determines that there is a configuration mismatch between the two ports to be linked, the application sends an alert message that includes a description of the possible causes of the unexpected port interruption on the first network device. In some examples, the description of the possible causes of the unexpected port interruption may include a description of the port configuration mismatch between the ports.

[0095] Accordingly, based on the described or indicated mismatch, a network administrator or other user can attempt to correct the mismatch (e.g., by configuring each port with a matching or compatible value for a previously mismatched setting) to enable communication between the ports. After identifying and displaying the cause of the link interruption (e.g., including a setting mismatch) as part of operation 240, flowchart 200 can return to operation 208, and the application at the first network device can execute the operations of flowchart 200 that began in operation 208 to identify the cause of the link interruption for the next unexpectedly interrupted port in the list identified from operation 204.

[0096] Although the triggers in the application at the first network device used to establish a peer-to-peer session (e.g., a TCP session or communication) with the second network device are in Figure 2In the example, operation 204 identifies the port that was unexpectedly interrupted, but the application can respond to any query on that specific port regardless of its state (e.g., connected or not connected). By having the application respond to any query on a specific port, the application can assist in identifying other causes of link failure or interruption, such as transceiver problems and loose connections (e.g., "peer unavailable," "peer unreachable," "cable / connection problem," etc.).

[0097] Figure 3 A method 300 is described, for example, that can be used to identify the causes of link outages as described herein. For example, method 300 can be used to determine the cause of an outage at a first port of a first network device or the cause of a link outage between a first port and a second port of a second network device (e.g., a link or port failure). Accordingly, using the techniques described herein, information maintained by a topology manager (e.g., a topology file or map maintained by the PTM) can be used to exchange link-based parameters between the first and second network devices to find possible causes of the first port outage and / or the link outage between the first and second ports.

[0098] Method 300 (and / or one or more of its steps) may be performed, for example, by at least one processor or otherwise. The at least one processor may be the same as or similar to the processor described above. The at least one processor may be part of a computing device (such as a personal computer, laptop, smartphone, etc.) or part of the network device or peer-to-peer device described above. Any processor other than those described herein may also be used to perform method 300. The at least one processor may perform method 300 by executing elements stored in memory such as that of the computing device, network device, or peer-to-peer device described above. Elements stored in memory and executed by the processor may cause the processor to perform one or more steps of the functions shown in method 300. One or more portions of method 300 may be executed by the processor executing any contents of memory.

[0099] Method 300 includes determining that a first port of the first peer device has been unexpectedly changed to a port interrupted state (step 304). For example, a list of the ports initialized and configured on the first peer device can be obtained to identify which ports were expected to be connected but are interrupted. In some examples, ports expected to be connected but interrupted can be identified based on ports that are managed to be configured to be connected.

[0100] Method 300 further includes referencing a topology file to determine a second port of the second peer device, to which the first peer device intends to establish a link (step 308) if the first port is not in a port interruption state. For example, the topology file may be maintained by a topology manager or PTM and may include information about the ports to be linked. In some examples, the first peer device may include a first LFA or a first LFA application, while the second peer device may include a second LFA or a second LFA application, wherein the first and second LFAs reference the topology file via PTM. Alternatively, a single LFA or LFA application may be provided in either the first or second peer device, wherein the single LFA application references the topology file via PTM.

[0101] Method 300 further includes determining that the port settings of the first port do not match the associated port settings of the second port (step 312). In some examples, the port settings of the first port and the second port can be determined to be mismatched based on LFA exchange information for each network device (e.g., via a TCP session). For example, a first LFA provided on a first peer device can communicate with a second LFA provided on a second peer device via the management network to exchange port settings. To enable these communications, the LFA provided at the first peer device (e.g., the first of two LFAs or a single LFA) can be configured to obtain hostname and interface name information (and any other necessary information) corresponding to the second port and / or the second peer device from a topology file.

[0102] Method 300 further includes sending an alert message to a network administrator in response to determining that the port settings of the first port do not match the associated port settings of the second port (step 316). For example, the alert message may include a description of possible causes for the port interruption state. In some examples, the description of possible causes for the port interruption state may include a description of a port setting mismatch between the first and second ports.

[0103] This disclosure includes embodiments of method 300, which include more or fewer steps than those described above, and / or one or more steps that are different from those described above.

[0104] Figure 4 A method 400 is described that can be used, for example, to identify link outage causes that are not the result of configuration mismatch as described herein.

[0105] Method 400 (and / or one or more of its steps) may be implemented, for example, by at least one processor or otherwise performed. The at least one processor may be the same as or similar to the processor described above. The at least one processor may be part of a computing device (such as a personal computer, laptop, smartphone, etc.) or part of the network device or peer-to-peer device described above. Any processor other than those described herein may also be used to perform method 400. The at least one processor may perform method 400 by executing elements stored in memory such as the memory of the computing device, network device, or peer-to-peer device described above. Elements stored in memory and executed by the processor may cause the processor to perform one or more steps of the functions shown in method 400. One or more portions of method 400 may be executed by the processor executing any contents of memory.

[0106] The method 400 includes determining that a first port of the first peer device has unexpectedly changed to a port outage state (step 404). The method 400 also includes referring to a topology file to determine a second port of the second peer device, to which the first peer device intends to establish a link if the first port is not in a port outage state (step 408).

[0107] Method 400 further includes determining (e.g., at an LFA or LFA application provided at one or both of the first and second peer devices) that the port settings of the first port match the port settings of the second port (step 412). For example, the LFA provided at the first peer device can determine the port settings of the second port (e.g., through a TCP session with a second LFA at the second peer device, through a topology file, etc.), and these port settings of the second port can be cross-checked with the port settings of the first peer device to determine any mismatches or absences in settings.

[0108] Method 400 further includes: including information describing the cable or connection problem as a possible source of the link failure (e.g., based on determining that the port settings of the first port match the port settings of the second port) (step 416). For example, if the port settings of the two ports to be linked match (or are compatible), but the first port is still in a port outage state and / or there is a link failure between the ports, then the possible source of the link failure can be determined to be a cable or connection problem. This cable or connection problem can include or indicate that the connection between the ports is incorrect, the port configurations and settings are compatible, but a problem with the connection or cable causes the link failure or outage, a configuration mismatch is identified and corrected, but the link is still outage, the connection and configuration are correct and compatible, but the link is still outage, or a combination thereof. In such an example, an alert message can be sent to the network administrator (or other user) including a description of the possible causes of the link failure.

[0109] This disclosure includes embodiments of method 400, which include more or fewer steps than those described above, and / or one or more steps that are different from those described above.

[0110] Figure 5 A method 500 is described that can be used, for example, to implement communication between applications that provide communication between different peer devices to compare port settings as described herein on each peer device.

[0111] Method 500 (and / or one or more steps thereof) may be executed, for example, by at least one processor or otherwise. The at least one processor may be the same as or similar to the processor described above. The at least one processor may be part of a computing device (such as a personal computer, laptop, smartphone, etc.) or part of the network device or peer-to-peer device described above. Any processor other than those described herein may also be used to execute method 500. The at least one processor may execute method 500 by executing elements stored in memory such as the memory of the computing device, network device, or peer-to-peer device described above. Elements stored in memory and executed by the processor may cause the processor to perform one or more steps of the functions shown in method 500. One or more portions of method 500 may be executed by the processor executing any contents of memory.

[0112] Method 500 includes determining that a first port of the first peer device has unexpectedly changed to a port interruption state (step 504). Method 500 also includes referring to a topology file to identify a second port of the second peer device, to which the first peer device intends to establish a link if the first port is not in a port interruption state (step 508). In some examples, a first LFA (or a first LFA application) may be provided at the first peer device, and a second LFA (or a second LFA application) may be provided at the second peer device.

[0113] Method 500 further includes enabling the first LFA and the second LFA to exchange port settings of the first port and the second port (step 512). For example, the first LFA can communicate with the second LFA via a management network (e.g., using a TCP session) to exchange port settings.

[0114] Method 500 further includes configuring a first LFA to compare the port settings of a first port with the port settings of a second port to determine whether the port settings of the first port match the port settings of the second port (step 516). Method 500 also includes determining whether the port settings of the first port match the port settings of the second port (step 520). In some examples, if the port settings of the first port match the port settings of the second port, but the first port is still in a port interruption state and / or the link between the first and second ports is still interrupted, the cause of the port interruption state can be determined as a connection or cable problem and displayed accordingly. Alternatively, if the port settings of the first port and the port settings of the second port do not match, the corresponding mismatch can be logged, and an alarm message indicating the mismatch can be displayed.

[0115] This disclosure includes embodiments of method 500, which include more or fewer steps than those described above, and / or one or more steps that are different from those described above.

[0116] As stated above, this disclosure covers those with less than Figure 3 , Figure 4 and Figure 5 Methods that encompass all steps defined in (and the corresponding descriptions of methods 300, 400, and 500), as well as methods covering beyond... Figure 3 , Figure 4 and Figure 5 (and the corresponding descriptions of methods 300, 400, and 500) of the additional steps identified herein. This disclosure also covers methods that include one or more steps from one method described herein, and methods that include one or more steps from another method described herein. Any association described herein may be or includes registration or any other association.

[0117] Any steps, functions, and operations discussed in this article can be performed continuously and automatically.

[0118] Exemplary systems and methods of this disclosure have been described in conjunction with dual-connection switching modules. However, to avoid unnecessarily obscuring the scope of this disclosure, some known structures and apparatuses have been omitted from the foregoing description. Such omissions should not be construed as limiting the scope of the claimed disclosure. Specific details are set forth in order to provide an understanding of this disclosure. However, it should be understood that this disclosure may be implemented in various ways beyond the specific details set forth herein.

[0119] Some variations and modifications to this disclosure may be used. It is possible to provide certain features of this disclosure without providing others.

[0120] The terms "an embodiment," "embodiment," "exemplary embodiment," "some embodiments," etc., used in this specification indicate that the described embodiment may include a specific feature, structure, or characteristic, but each embodiment does not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Additionally, when a specific feature, structure, or characteristic is described in connection with an embodiment, the description of that feature, structure, or characteristic is applicable to any other embodiment, unless so stated and / or readily apparent to those skilled in the art from the description. This disclosure includes, in various embodiments, configurations, and aspects, components, methods, processes, systems, and / or apparatuses substantially as described herein, including various embodiments, sub-combinations, and subsets thereof. Those skilled in the art, upon understanding this disclosure, will understand how to make and use the systems and methods disclosed herein. This disclosure includes, in various embodiments, configurations, and aspects, providing apparatus and processes in the absence of items not depicted and / or described herein, or in various embodiments, configurations, or aspects herein, including the absence of such items that could potentially be used in prior equipment or processes, for example, to improve performance, ease of implementation, and / or reduce implementation costs.

[0121] The foregoing discussion of the disclosure is for illustrative and descriptive purposes. The foregoing is not intended to limit the disclosure to the form or format disclosed herein. In the foregoing detailed description, various features of the disclosure have been combined in one or more embodiments, configurations, or aspects for the purpose of simplification. Features of embodiments, configurations, or aspects of the disclosure may be combined in other embodiments, configurations, or aspects beyond those discussed above. This approach to disclosure should not be construed as reflecting an intention that the claimed disclosure requires more features than expressly recited in each claim. Rather, as reflected in the following claims, the inventive aspect lies in fewer than all features of a single embodiment, configuration, or aspect of the foregoing disclosure. Therefore, the following claims are hereby incorporated into this detailed description, each claim existing independently as a separate preferred embodiment of the disclosure.

[0122] Furthermore, although the description of this disclosure has included descriptions of one or more embodiments, configurations, or aspects, as well as certain variations and modifications, other variations, combinations, and modifications are also within the scope of this disclosure, for example, which may fall within the technical and knowledge scope of those skilled in the art after understanding the content of this disclosure. The rights to be claimed include alternative embodiments, configurations, or aspects within the permissible scope, including alternating, interchangeable, and / or equivalent structures, functions, scopes, or steps to the claimed structure, function, scope, or steps, whether such alternating, interchangeable, and / or equivalent structures, functions, scopes, or steps are disclosed herein, and no patentable subject matter is intended to be disclosed.

[0123] As used herein, the singular forms “a,” “an,” and “the” are intended to also include the plural forms, unless the context explicitly indicates otherwise. It will be further understood that the terms “include,” “including,” “includes,” “comprise,” “comprises,” and / or “comprising,” when used in this specification, specify the presence of the stated feature, integer, step, operation, element, and / or component, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. The terms “and / or” include any and all combinations of one or more associated listed items.

[0124] As used herein, the term "automatic" and its variations refer to any process or operation that is typically continuous or semi-continuous and does not require substantial human input in its execution. However, a process or operation can be automatic even if its execution uses substantial or non-substantial human input, if the input is received prior to the execution of the process or operation. Human input is considered substantial if it influences how a process or operation is executed. Human input consenting to the execution of the process or operation is not considered "substantial."

[0125] It should be understood that each maximum number limit given in this disclosure is considered to include each lower number limit, as alternatively as such lower number limit is explicitly stated herein. Each minimum number limit given in this disclosure is considered to include each higher number limit, as alternatively as such higher number limits are explicitly stated herein. Each number range given in this disclosure is considered to include each narrower number range belonging to such a wider number range, as if such narrower number ranges were explicitly stated herein.

Claims

1. A method comprising: It was determined that the first port of the first peer device had unexpectedly entered a port interrupt state; Referring to the topology file, the second port of the second peer device is identified. If the first peer device is not in a port interruption state, the first peer device is intended to have a link with the second peer device, wherein the first peer device includes a first link failure analyzer (LFA), wherein the second peer device includes a second LFA, and wherein the first LFA communicates with the second LFA through the management network to exchange port settings. It is determined that the port settings of the first port do not match the associated port settings of the second port; and In response to the determination that the port settings of the first port do not match the associated port settings of the second port, an alert message is sent to the network administrator.

2. The method of claim 1, wherein the alarm message includes a description of possible causes for the port interruption state.

3. The method of claim 2, wherein the description of possible causes of the port interruption state includes a description of a port setting mismatch between the first port and the second port.

4. The method of claim 1, wherein the topology file is maintained by a specified topology manager (PTM), wherein the first LFA references the topology file through the PTM, and wherein the second LFA also references the topology file through the PTM.

5. A system comprising: processor; as well as A memory coupled to and readable by the processor, wherein instructions are stored, when executed by the processor, cause the processor to: It was determined that the first port of the first peer device had been unexpectedly changed to a port interrupt state; Referring to the topology file, the second port of the second peer device is identified. If the first peer device is not in a port interruption state, the first peer device is intended to have a link with the second peer device, wherein the first peer device includes a first link failure analyzer (LFA), wherein the second peer device includes a second LFA, and wherein the first LFA communicates with the second LFA through the management network to exchange port settings. It is determined that the port settings of the first port do not match the associated port settings of the second port; and In response to the determination that the port settings of the first port do not match the associated port settings of the second port, an alert message is sent to the network administrator.

6. The system of claim 5, wherein the alarm message includes a description of possible causes for the port interruption state.

7. The system of claim 6, wherein the description of possible causes of the port interruption state includes a description of a port setting mismatch between the first port and the second port.

8. The system of claim 5, wherein the topology file is maintained by a specified topology manager (PTM), wherein the first LFA references the topology file through the PTM, and wherein the second LFA also references the topology file through the PTM.

9. A method comprising: It was determined that the first port of the first peer device had unexpectedly entered a port interrupt state; Referring to a topology file to identify a second port of the second peer device, if the first peer device intends to have a link with the second peer device, unless the first port is in a port interruption state, wherein a Link Failure Analyzer (LFA) is provided at one of the first peer device and the second peer device, wherein the topology file is maintained by a specified topology manager (PTM), and wherein the LFA refers to the topology file through the PTM; It is determined that the port settings of the first port do not match the associated port settings of the second port; and In response to the determination that the port settings of the first port do not match the associated port settings of the second port, an alert message is sent to the network administrator.

10. The method of claim 9, wherein the LFA is configured to obtain hostname and interface name information from the topology file.

11. The method of claim 9, further comprising: At the LFA, it is determined that the port settings of the first port match the port settings of the second port. as well as The alarm message includes information describing the cable or connection problem as a possible source of link failure.

12. The method of claim 11, wherein the LFA includes a first LFA provided at the first peer device, and wherein a second LFA is provided at the second peer device, the method further comprising: This enables the first LFA and the second LFA to exchange the port settings of the first port and the port settings of the second port. as well as Configure the first LFA to compare the port settings of the first port with the port settings of the second port to determine that the port settings of the first port match the port settings of the second port.

13. The method of claim 11, wherein the alarm message includes a description of possible causes of the link failure.

14. A system comprising: processor; as well as A memory coupled to and readable by the processor, wherein instructions are stored, when executed by the processor, cause the processor to: It was determined that the first port of the first peer device had been unexpectedly changed to a port interrupt state; Referring to a topology file to identify a second port of the second peer device, if the first peer device intends to have a link with the second peer device, unless the first port is in a port interruption state, wherein a Link Failure Analyzer (LFA) is provided at one of the first peer device and the second peer device, wherein the topology file is maintained by a specified topology manager (PTM), and wherein the LFA refers to the topology file through the PTM; It is determined that the port settings of the first port do not match the associated port settings of the second port; and In response to the determination that the port settings of the first port do not match the associated port settings of the second port, an alert message is sent to the network administrator.

15. The system of claim 14, wherein the LFA is configured to obtain host name and interface name information from the topology file.

16. The system of claim 14, wherein the instructions further cause the processor to: At the LFA, it is determined that the port settings of the first port match the port settings of the second port; and The alarm message includes information describing the cable or connection problem as a possible source of link failure.

17. The system of claim 16, wherein the LFA includes a first LFA provided at the first peer device, and wherein a second LFA is provided at the second peer device, the instructions further causing the processor to: Enables the first LFA and the second LFA to exchange the port settings of the first port and the port settings of the second port; and Configure the first LFA to compare the port settings of the first port with the port settings of the second port to determine that the port settings of the first port match the port settings of the second port.