Fault detection method and device

By carrying the statistical information of the device in the detection message and using the statistical information of the detection message of the first and second devices, the problem of difficulty in distinguishing the fault of the peer-end device from the intermediate link in the prior art is solved, and the fault positioning and system reliability are achieved.

CN120455414APending Publication Date: 2025-08-08HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410173588.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing inter-device heartbeat detection mechanism is difficult to distinguish between peer equipment failures and intermediate link failures, resulting in inaccurate fault detection.

Method used

By carrying the statistical information of the device in the detection message, using the statistical information of the detection message of the first and second devices, the fault location information is determined, including the fault of the local device, the fault of the opposite device or the intermediate link failure.

Benefits of technology

It realizes accurate positioning of peer equipment failures and intermediate link failures, improving the accuracy of fault detection and system reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455414A_ABST
    Figure CN120455414A_ABST
Patent Text Reader

Abstract

The invention provides a fault detection method and equipment, relates to the technical field of communication, and is used for accurately determining the position of a fault and further distinguishing an opposite-end equipment fault and an intermediate link fault. The method comprises the steps that first equipment obtains first equipment detection message statistical information and second equipment detection message statistical information, the first equipment detection message statistical information is statistical information of a detection message in a counter of the first equipment, and the second equipment detection message statistical information is statistical information of the detection message in a counter of the second equipment; the detection message is a detection request or a detection response; and under the condition that transmission between the first equipment and the second equipment fails, the first equipment determines fault position information according to the first equipment detection message statistical information and the second equipment detection message statistical information, and the fault position information comprises a home terminal equipment fault, an opposite terminal equipment fault or an intermediate link fault.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of communication technology, and in particular to a fault detection method and device. Background Art

[0002] In the use of high-speed switching chips, a guaranteed forwarding channel fault detection mechanism is required. For cluster computing systems, an effective fault detection mechanism is currently an important work direction for improving the reliability of cluster computing systems.

[0003] For heartbeat detection between devices, the existing forwarding channel fault detection mechanism mainly sends detection messages to other devices, and then determines whether there is a fault through the central processor by receiving the heartbeat messages responded by other devices.

[0004] In the current heartbeat detection mechanism between devices, for both peer device failure and intermediate link failure, the local device cannot receive the heartbeat message response from other devices, making it difficult for the local device to distinguish between peer device failure and intermediate link failure. Summary of the Invention

[0005] The embodiments of the present application provide a fault detection device that can accurately determine the location of a fault and further distinguish between a fault in an opposite-end device and a fault in an intermediate link.

[0006] In a first aspect, an embodiment of the present application provides a fault detection method, the method comprising: a first device obtains statistical information of a first device detection message and a second device detection message, the first device detection message statistical information being statistical information of a detection message counter of the first device, the second device detection message statistical information being statistical information of a detection message counter of the second device, and the detection message being a detection request or a detection response; in the event of a transmission failure between the first device and the second device, the first device determines fault location information based on the first device detection message statistical information and the second device detection message statistical information, the fault location information comprising: local device failure, opposite device failure, or intermediate link failure.

[0007] In this possible implementation, the first device can obtain detection message statistics of the first device and the second device, and the first device determines specific fault location information based on the detection message statistics of the first device and the second device, and can distinguish between a fault in the opposite device and a fault in the intermediate link.

[0008] In one possible implementation, before the first device obtains the detection message statistical information of the first device and the detection message statistical information of the second device, the above method also includes: the first device sends a first detection request to the second device, and the first detection request carries the detection message statistical information of the first device; the first device receives a first detection response from the second device, and the first detection response carries the detection message statistical information of the second device.

[0009] In this possible implementation, the detection messages sent by the first device and the second device (i.e., the first detection request and the first detection response) both carry corresponding detection message statistical information, so that the device receiving the detection message can obtain the statistical information of the counters of other devices, thereby knowing the forwarding status of the detection device in other devices.

[0010] In one possible implementation, there are multiple communication links between the first device and the second device, and the first device sending a first detection request to the second device includes: the first device sending the first detection request to the second device through the first link; the first device receiving a first detection response from the second device includes: the first device receiving the first detection response from the second device through the first link; the first device obtaining the first device detection message statistical information and the second device detection message statistical information includes: the first device obtaining the first device detection message statistical information and the second device detection message statistical information through the second link, the first link and the second link are included in multiple communication links, the first link is a faulty link, and the first link is a normal link.

[0011] In this possible implementation, the detection request is sent and the detection response is received through the first link, and the detection message statistical information of the first device and the detection message statistical information of the second device are obtained through the second link. Therefore, when the first link fails, resulting in the first device and the second device failing, the first device can obtain the first device detection message statistical information and the second device detection message statistical information through the normally working second link, so that the failure of the first link still does not affect the first device's judgment of the location of the first link failure.

[0012] In a possible implementation, the detection message includes a detection identifier, a type identifier, a message source address, and a message destination address, wherein the detection identifier is used to indicate that the detection message is a detection message, and the type identifier is used to indicate that the detection message is a detection request or a detection response.

[0013] In a possible implementation, the detection message further includes status information of multiple interfaces, where the status information indicates whether the corresponding interfaces are in an up state or a down state.

[0014] In one possible implementation, the first device obtains the first device detection message statistical information and the second device detection message statistical information, including: the first device obtains the first device detection message statistical information through the second detection request; the first device obtains the second device detection message statistical information through the second detection response, and the second detection request and the second detection response are detection messages corresponding to the second link.

[0015] In this possible implementation, the first device obtains the detection message statistical information through the detection message corresponding to the second link, so that the failure of the first link still does not affect the first device's acquisition of the detection message statistical information.

[0016] In one possible implementation, the first device determines the fault location information based on the first device detection message statistical information and the second device detection message statistical information, including: the first device determines the transmission interruption point of the detection message based on the first device detection message statistical information and the second device detection message statistical information; the first device determines the fault location information based on the transmission interruption point.

[0017] In this possible implementation, the first device determines a transmission interruption point through detection message statistical information, where the interruption point indicates a location where a fault occurs in the detection message.

[0018] In a possible implementation, the method further includes: the first device determining that the fault is a unidirectional fault or a bidirectional fault.

[0019] In this possible implementation, the first device may also determine whether the fault is a unidirectional fault or a bidirectional fault based on detection message statistical information, thereby enhancing the fault detection capability.

[0020] In a possible implementation, the method further includes: the first device performing fault processing according to the fault location information.

[0021] In one possible implementation, the first device performs fault processing based on the fault location information, including: if the fault location information is a fault in the opposite device, the first device reports an alarm or does not perform any operation; or if the fault location information is a fault in the intermediate link, the first device reports an alarm; or if the fault location information is a fault in the local device, the first device reports an alarm or performs a fault reset operation.

[0022] In a possible implementation, the local device failure includes a unidirectional failure of the local device and a bidirectional failure of the local device; the opposite device failure includes a unidirectional failure of the opposite device and a bidirectional failure of the opposite device; and the intermediate link failure includes a unidirectional failure of the intermediate link and a bidirectional failure of the intermediate link.

[0023] In second aspect, an embodiment of the present application provides a fault detection method, which includes: a second device receives a first detection request from a first device, the first detection request carries detection message statistical information of the first device; the second device sends a first detection response to the first device, the first detection response carries detection message statistical information of the second device.

[0024] In a third aspect, embodiments of the present application provide a first device comprising: a processor and a memory. The processor is coupled to the memory; the memory is configured to store computer instructions, which are loaded and executed by the processor to cause the first device to implement any one of the methods provided in the first, second, or fifth aspects.

[0025] In a fourth aspect, embodiments of the present application provide a second device comprising: a processor and a memory. The processor is coupled to the memory; the memory is configured to store computer instructions, which are loaded and executed by the processor to cause the second device to implement any one of the methods provided in the first, third, or fourth aspects.

[0026] In a fifth aspect, an embodiment of the present application provides a chip comprising: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; the processor is used to run the code instructions to execute any one of the methods provided in the first aspect, the second aspect or the fifth aspect.

[0027] In a sixth aspect, an embodiment of the present application provides a chip comprising: a processor and an interface circuit; the interface circuit is used to receive code instructions and transmit them to the processor; and the processor is used to run the code instructions to execute any one of the methods provided in the first aspect, the third aspect, or the fourth aspect.

[0028] In the seventh aspect, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer program instruction, and the computer program instruction is loaded and executed by a processor to implement any one of the methods provided in the first, second or fifth aspects above.

[0029] In an eighth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores at least one computer program instruction, and the computer program instruction is loaded and executed by a processor to implement any one of the methods provided in the first, third or fourth aspects above.

[0030] In a ninth aspect, an embodiment of the present application provides a computer program product, comprising computer execution instructions, which, when the computer execution instructions are run on a computer, enable the computer to execute any one of the methods provided in the first aspect, the second aspect, or the fifth aspect.

[0031] In the tenth aspect, an embodiment of the present application provides a computer program product, including computer execution instructions, which, when the computer execution instructions are run on a computer, enable the computer to execute any one of the methods provided in the first aspect, the third aspect or the fourth aspect.

[0032] The technical effects brought about by any implementation method in the third aspect to the tenth aspect can be referred to the technical effects brought about by the corresponding implementation methods in the first aspect to the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] Figure 1a A schematic diagram of a forwarding channel fault detection mechanism.

[0034] Figure 1b This is a scenario diagram of another forwarding channel failure detection mechanism;

[0035] Figure 2 A schematic diagram of a scenario of a fault detection method provided in an embodiment of the present application;

[0036] Figure 3 A flowchart of a fault detection method provided in an embodiment of the present application;

[0037] Figure 4 A schematic diagram of a detection message provided in an embodiment of the present application;

[0038] Figure 5 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0039] Figure 6 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0040] Figure 7 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0041] Figure 8 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0042] Figure 9 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0043] Figure 10 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0044] Figure 11 A schematic diagram of another fault detection method provided in an embodiment of the present application;

[0045] Figure 12A schematic structural diagram of a first device provided in an embodiment of the present application;

[0046] Figure 13 A schematic structural diagram of a second device provided in an embodiment of the present application;

[0047] Figure 14 A schematic structural diagram of another first device provided in an embodiment of the present application;

[0048] Figure 15 A schematic structural diagram of another second device provided in an embodiment of the present application;

[0049] Figure 16 A schematic diagram of the structure of a communication system provided in an embodiment of the present application. DETAILED DESCRIPTION

[0050] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.

[0051] In the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0052] In addition, to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that the words "first" and "second" do not limit the quantity or execution order, and the words "first" and "second" do not necessarily mean different.

[0053] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be construed as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner to facilitate understanding.

[0054] It will be understood that the “embodiment” mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the various embodiments throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It will be understood that in the various embodiments of the present application, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present application.

[0055] It is understood that some optional features in the embodiments of the present application may, in certain scenarios, be implemented independently of other features, such as the solution on which they are currently based, to solve corresponding technical problems and achieve corresponding effects. They may also be combined with other features in certain scenarios as needed. Accordingly, the devices provided in the embodiments of the present application may also implement these features or functions accordingly, which will not be described in detail here.

[0056] In this application, unless otherwise specified, the same or similar parts between the various embodiments can be referenced to each other. In this application, unless otherwise specified and there is no logical conflict between the various embodiments, the terms and / or descriptions between different embodiments are consistent and can be referenced to each other. Different embodiments can be combined to form new embodiments based on their inherent logical relationships. The following implementation methods of this application do not constitute a limitation on the scope of protection of this application.

[0057] High-performance computing (HPC) and artificial intelligence (AI) cluster computing systems utilize a large number of processors and switching nodes. The complex system architecture presents challenges in system availability. For example, high-speed switch chips may fail during operation. Effectively preventing these failures is a key area of focus for system reliability. While high-speed switch chips have comprehensive Failure Mode and Effects Analysis (FMEA) capabilities, limited by their area resources and implementation complexity, FMEA mechanisms cannot be added indefinitely. Therefore, a robust forwarding path fault detection mechanism is necessary for high-speed switch chip deployment scenarios. Furthermore, highly integrated systems like cluster computing require an effective system-level silent fault detection mechanism to improve the forwarding path fault detection coverage of the high-speed switch chips, thereby enhancing system reliability. In summary, effective fault detection mechanisms are currently a key area of focus for improving cluster computing system reliability.

[0058] like Figure 1a and Figure 1b As shown, the existing forwarding channel fault detection mechanism mainly includes two fault detection mechanisms: intra-device loopback heartbeat detection and inter-device heartbeat detection. Among them, the intra-device loopback heartbeat detection constructs a timed heartbeat message through the control plane, and then inserts the timed heartbeat message into the chip forwarding plane. After the forwarding plane loops back and sends it up, the control plane central processing unit (CPU) determines whether there is a fault. Figure 1a In the example, the CPU of switch 1 regularly generates a heartbeat message and sends it to the forwarding chip, which then loops back and forwards the heartbeat message to the CPU of switch 1. The CPU can determine whether there is a fault in switch 1 based on the received heartbeat message.

[0059] like Figure 1b As shown in , two adjacent devices can detect forwarding channel failures through inter-device heartbeat detection, and the two devices independently send request messages and response messages. Specifically, Figure 1b As shown, if there is a fault point in the forwarding channel formed by the CPU1 of switch 1, the forwarding chip 1 of switch 1, the forwarding chip 2 of switch 2 and the CPU2 of switch 2, it will affect the transmission of request messages and response messages of switch 1 and switch 2. Switch 1 and switch 2 can determine whether there is a faulty node in the forwarding channel based on this.

[0060] However, in the current inter-device heartbeat detection mechanism, it is difficult for the local device to distinguish between the peer device failure that does not require alarm reporting and the intermediate link failure that requires alarm reporting. Figure 1b In the example, counting point 4 is located at CPU 1 of switch 1, counting point 3 is located at forwarding chip 1 of switch 1, counting point 2 is located at forwarding chip 2 of switch 2, and counting point 1 is located at CPU 2 of switch 2.

[0061] For example, after a fault occurs at fault point 1, for switch 2, counting point 2 does not receive the heartbeat message from switch 1, and switch 2 cannot determine whether it is a local chip fault, a peer chip fault, or an intermediate link fault; for switch 1, switch 1 also does not receive the heartbeat message from switch 2, and switch 1 cannot determine whether it is a local chip fault, a peer chip fault, or an intermediate link fault.

[0062] For example, after a fault occurs at fault point 2, for switch 2, counting point 2 receives a heartbeat message from switch 1, but counting point 1 does not receive a heartbeat message from switch 1. Therefore, switch 2 can determine that it is an internal chip fault; however, for switch 1, switch 1 does not receive a heartbeat message from switch 2, so switch 1 cannot determine whether it is a local chip fault, a peer chip fault, or an intermediate link fault.

[0063] In summary, in the current heartbeat detection mechanism between devices, it is difficult for a device to distinguish between a peer device failure and an intermediate link failure.

[0064] An embodiment of the present application provides a fault detection method that can be used to distinguish between a peer device fault and an intermediate link fault. The method includes: a first device obtains first device detection message statistical information and second device detection message statistical information, wherein the first device detection message statistical information is statistical information of the detection message counter of the first device, and the second device detection message statistical information is statistical information of the detection message counter of the second device, and the detection message is a detection request or a detection response; in the event of a transmission fault between the first device and the second device, the first device determines fault location information based on the first device detection message statistical information and the second device detection message statistical information, and the fault location information includes: local device fault, peer device fault, or intermediate link fault.

[0065] like Figure 2 As shown, in an embodiment of the present application, for two adjacent switching devices connected by multiple links, the working status of the forwarding plane can be detected by using two detection messages, detection request and detection response, that is, determining whether there is a transmission fault and the specific location of the transmission fault point. Specifically, the device can use the message flow count in the system to count the detection messages. When the device sends a detection message, it will carry the statistical counting information of the detection message in the device in the detection message. Then, the two devices can comprehensively determine the location of the fault point in the link through the statistical information of the detection message of the device and the statistical information of the detection message of the neighboring device.

[0066] When multiple paths exist between two devices, each device can include full or multi-path detection statistics between the two devices in each link's detection message when sending a detection message. This allows the peer device to obtain statistics for detection messages on the broken link even if a link fails, thereby enhancing fault detection adequacy and increasing the feasibility of implementing the embodiments of the present application.

[0067] In the embodiment of the present application, the communication method between the first device and the second device can be various, and the communication method can be wired communication or wireless communication, which is not specifically limited here.

[0068] In the embodiments of the present application, the first device and the second device are not limited to specific device types, such as operator routers, campus switches, data center switches and other data communication equipment. As long as the neighboring node needs to send a detection message to check whether the forwarding channel is normal, the fault detection method of the present application can be applied.

[0069] It is understood that in the embodiments of the present application, the execution subject may perform some or all of the steps in the embodiments of the present application. These steps or operations are merely examples, and the embodiments of the present application may also perform other operations or variations of various operations. In addition, the various steps may be performed in a different order than those presented in the embodiments of the present application, and it is possible that not all operations in the embodiments of the present application need to be performed.

[0070] It should be noted that the message names between the devices or the names of the parameters in the messages in the following embodiments of the present application are only examples. Other names may be used in specific implementations, and the embodiments of the present application do not specifically limit this.

[0071] like Figure 3 As shown, an embodiment of the present application provides a method for fault detection, the method comprising:

[0072] 301. Send a first detection request.

[0073] The first device sends a first detection request to the second device, where the first detection request includes detection message statistics information of the first device.

[0074] Correspondingly, the second device receives the first detection request from the first device.

[0075] In an embodiment of the present application, both the detection request and the corresponding detection response will have a characteristic identification bit to indicate that the message is a detection message, so that the device receiving the detection message can quickly and accurately determine that it is a detection message.

[0076] Specifically, for example, in a communication system between a first device and a second device, a bit is set in the message header of the message as a characteristic identification bit. If the characteristic identification bit is "1", it indicates that the message is a detection message; if the characteristic identification bit is "0", it indicates that the message is not a detection message. For example, when the first device sends a detection request to the second device, the characteristic identification bit of the first detection request is "1", indicating that the detection request is a detection message. In addition, the message can also be identified as a detection message in other ways, such as using the first bit of the payload data of the detection message as the characteristic identification bit, and the characteristic identification bit is "1", indicating that the message is a detection message; if the characteristic identification bit is "0", it indicates that the message is not a detection message. There can also be more implementation methods, such as indicating that the message is a detection message through other messages, which are not limited here.

[0077] In an embodiment of the present application, the detection message may further include a detection message type identification bit to indicate the type of the detection message, that is, to indicate that the message is a detection request or to indicate that the message is a detection response.

[0078] Specifically, for example, in the payload data of a detection message, the first 4 bytes serve as the detection message header. The lowest bit is used to distinguish between a detection request and a detection response, which is defined as the "detection message type identification bit." A value of 0 for this bit indicates that the detection message is a detection request, and a value of 1 indicates that the detection message is a detection response. Other bits can be reserved. In addition, the type of the detection message can also be indicated by other means, such as by using other messages to indicate the type of the detection message, or by using a detection message type identification bit set in the message header of the detection message to indicate the type of the detection message. The specifics are not limited here.

[0079] In an embodiment of the present application, the local node address (i.e., the network address of the first device) can be used as the source address of the detection request message, and the destination node address (i.e., the network address of the second device) can be used as the destination address of the detection request.

[0080] To sum up, by setting the corresponding characteristic identification bit in the detection message to indicate whether the message is a detection message, setting the corresponding detection message type identification bit to indicate the type of the detection message, and setting the message source address and destination address, the device can determine the type of the message, which can be the detection request sent by this device, the detection response sent by this device, the detection request sent by the opposite device, and the detection response sent by the opposite device.

[0081] In an embodiment of the present application, the detection message carries the detection message statistical information of the first device. This detection message statistical information can be obtained by the first device through various counters in the first device. For example, as shown in Table 1, the first device includes a link reception counter, a link transmission counter, a route forwarding reception counter, a route forwarding transmission counter, a host reception counter, a host transmission counter, a CPU reception counter, and a CPU transmission counter. The indexes of these counters are all interface indexes. Through these counters, the first device can obtain statistical information about the transmission of the detection message within the first device. The statistical information can also include a timestamp corresponding to the counting information.

[0082] Table 1

[0083] Counter Index Message counting point Interface Index Link Receive Counter Interface Index Link send counter Interface Index Routing forwarding reception counter Interface Index Routing forwarding send counter Interface Index Host receive counter Interface Index Host send counter Interface Index CPU Receive Counter Interface Index CPU send counter

[0084] In an embodiment of the present application, the detection message statistical information carried by the detection message is detection message statistical information including multiple links. For example, there are three links between the first device and the second device. The detection message statistical information carried by the detection message can be the first link and the second link, or the first link and the third link, or the second link and the third link, or the first link, the second link and the third link. The specific information is not limited here.

[0085] In addition, the detection message statistics in the embodiment of the present application may also include status information of each interface in the detection state and interface transceiver count information. The interface status information includes the up or down status of each interface in the detection state and the transceiver count of the interface recorded in the message.

[0086] In an embodiment of the present application, the detection message may also carry more information, such as increasing the buffer cache depth to achieve more functions, which is not limited here.

[0087] In one possible implementation, Figure 4 As shown, the detection message in the embodiment of the present application may include a heartbeat message header, the full statistical count of the node, the full interface status of all nodes, and the transceiver statistics of all interfaces of the node. The heartbeat message header is a 4-byte heartbeat message header; the full statistical count of the node includes the detection message transceiver count of each link and the corresponding timestamp. After receiving the detection message, the peer device will save the detection message transceiver count information with the latest timestamp.

[0088] 302. Receive a first detection response.

[0089] In response to the first detection request, the second device sends a first detection response to the first device. Correspondingly, the first device receives the first detection response from the second device, the first detection response carrying detection message statistics information of the second device.

[0090] In the embodiment of the present application, the detection response and the detection request are both detection messages. Therefore, the manner in which the detection response carries the detection message statistical information of the second device, and the content of the carried detection message statistical information are the same as the detection message statistical information carried by the detection request in step 301. For details, please see the relevant description of the detection message carrying the detection message statistical information in step 301, which will not be repeated here.

[0091] In the embodiment of this application,

[0092] 303. Obtain statistical information of detection messages of the first device and statistical information of detection messages of the second device.

[0093] The first device obtains the detection message statistical information of the first device and the detection message statistical information of the second device.

[0094] Specifically, the first device obtains detection message statistical information of the second device based on the detection message sent by the second device, where the detection message statistical information of the second device indicates statistical information of the detection message by the counting point of the second device, and can reflect the forwarding status of the detection message in the second device. The first device obtains the detection message statistical information of the first device, where the detection message statistical information of the first device indicates statistical information of the detection message by the counting point of the first device, and can reflect the forwarding status of the detection message in the second device.

[0095] In an embodiment of the present application, for the first detection message (first detection request and first detection response) transmitted through the first link, after the first link fails, the detection message can carry detection message statistical information of multiple links, so the first device can obtain the statistical information of the first detection message through the normal second link, for example, through the second detection message (second detection request and second detection response) transmitted by the second detection message.

[0096] In one possible implementation, as shown in Table 2, for example, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request 1 to switch 2 every second, and switch 2 responds with a detection response 1 to switch 1. After 10 cycles, a fault occurs between switch 1 and switch 2. After five detection cycles, as shown in Table 2, switch 1 can determine, based on detection request 2 and detection response 2, detection message statistics of switch 1 for detection request 1 and detection response 1, and detection message statistics of switch 2 for the first detection request 1 and the first detection response 1.

[0097] In an embodiment of the present application, there are multiple links between the first device and the second device. After the link transmitting the detection request 1 and the detection response 1 fails, the first device and the second device can communicate through other redundant links. Therefore, after the link 1 transmitting the detection request 1 and the detection response 1 fails, the first device can still obtain the detection message statistical information of the second device through the detection request 2 and the detection response 2. The transmission link corresponding to the detection request 2 and the detection response 2 is link 2.

[0098] Table 2

[0099]

[0100] 304. Determine the fault location information.

[0101] The first device determines fault location information according to the first device detection message statistics information and the second device detection message statistics information, where the fault location information includes a local device fault, a peer device fault, or an intermediate link fault.

[0102] Specifically, the first device determines the transmission interruption point of the detection message according to the first device detection message statistical information and the second device detection message statistical information, and then can determine the specific fault location and the type of the fault according to the transmission terminal of the detection message.

[0103] In the embodiment of the present application, the first device determines a counting result statistical matrix based on the bidirectional counting results of the detection message counter, i.e., the first device detection message statistical information and the second device detection message statistical information. The counting result statistical matrix can be in the form of a table such as Table 2. The first device can locate the fault point based on the counting result statistical matrix and the forwarding process of the detection message, and determine whether it is a local device fault, a peer device fault, or an intermediate link fault.

[0104] In an embodiment of the present application, the counting result statistical matrix includes statistical information of the two forwarding processes of detection request and detection response. The first device can comprehensively determine the location information of the fault point from these two directions. If an interruption point occurs in only one direction, it is a unidirectional fault; if an interruption point occurs in both directions, it is a bidirectional fault.

[0105] Example 1: A unidirectional fault occurs on the peer device.

[0106] In one possible implementation, Figure 5 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 3.

[0107] Table 3

[0108]

[0109] As shown in Table 3, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 to 5, and a count of 10 between counters 6 and 8. Therefore, it can be determined that a transmission failure occurred between counters 5 and 6 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 5 to counter 6.

[0110] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 can be transmitted normally. However, the detection response sent by switch 1 to switch 2 has a count of 15 in counters 1 to 5 and a count of 10 between counters 6 and 8. Therefore, it can be determined that the transmission direction from switch 1 to switch 2 is a transmission failure between counters 5 and 6. Therefore, the detection request from the 11th to the 15th switching cycle cannot be transmitted from counter 5 to counter 6.

[0111] In summary, after the 10th switching cycle, a fault occurred between counters 5 and 6 in the test message transmission from Switch 1 to Switch 2, while Switch 2 was transmitting the test message normally to Switch 1. Therefore, it can be determined that a unidirectional transmission fault occurred between counters 5 and 6. Because the area between counters 5 and 6 is internal to Switch 2, Switch 1 can determine that the fault is a fault on the peer device.

[0112] Example 2: A unidirectional failure occurs on an intermediate link.

[0113] In another possible implementation, Figure 6 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 4.

[0114] Table 4

[0115]

[0116] As shown in Table 4, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 and 4, and a count of 10 between counters 5 and 8. Therefore, it can be determined that a transmission failure occurred between counters 4 and 5 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 5 to counter 6.

[0117] After the 10th exchange cycle, the detection request sent by switch 2 to switch 1 can be transmitted normally. However, the detection response sent by switch 1 to switch 2 has a count of 15 between counters 1 to 4 and a count of 10 between counters 5 and 8. Therefore, it can be determined that the transmission direction from switch 1 to switch 2 is a transmission failure between counters 4 and 5. The detection requests from the 11th to 15th exchange cycles cannot be transmitted from counter 4 to counter 5.

[0118] In summary, after the tenth switching cycle, a failure occurred between counters 4 and 5 in the test message transmission from Switch 1 to Switch 2, while Switch 2 was transmitting the test message normally to Switch 1. Therefore, it can be determined that a unidirectional transmission failure occurred between counters 4 and 5. Because the link between counters 4 and 5 is an intermediate link between Switches 1 and 2, Switch 1 can determine that the fault is an intermediate link failure.

[0119] The intermediate link failure in the embodiment of the present application refers to a failure in which the fault point is between counting point 4 and counting point 5. The fault point is specifically within the first device, or within the second device, or on the optical link between the two. Because the two counting points are only used to transmit signals, they can be understood as an intermediate link, which is not specifically limited here.

[0120] Example 3: Bidirectional failure of the peer device.

[0121] In one possible implementation, Figure 7 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 3.

[0122] Table 5

[0123]

[0124] As shown in Table 5, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 to 5, and a count of 10 between counters 6 and 8. Therefore, it can be determined that a transmission failure occurred between counters 5 and 6 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 5 to counter 6.

[0125] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 has a count of 15 between counter 8 and counter 6, and a count of 10 between counter 5 and counter 1. Therefore, it can be determined that a transmission failure occurred between counter 5 and counter 6 in the direction of transmission from switch 2 to switch 1. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 6 to counter 5.

[0126] In summary, after the tenth switching cycle, a failure occurred between counters 5 and 6 in the test message transmission from Switch 1 to Switch 2, and a failure also occurred between counters 5 and 6 in the test message transmission from Switch 2 to Switch 1. Therefore, it can be determined that a bidirectional transmission failure occurred between counters 5 and 6. Because the area between counters 5 and 6 is internal to Switch 2, Switch 1 can determine that the fault is a fault in the peer device.

[0127] Example 4: Bidirectional failure of the intermediate link.

[0128] In one possible implementation, Figure 8 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 6.

[0129] Table 6

[0130]

[0131] As shown in Table 6, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 and 5, and a count of 10 between counters 6 and 8. Therefore, it can be determined that a transmission failure occurred between counters 4 and 5 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 4 to counter 5.

[0132] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 has a count of 15 between counter 8 and counter 5, and a count of 10 between counter 4 and counter 1. Therefore, it can be determined that a transmission failure occurred between counter 4 and counter 5 in the transmission direction of the message sent by switch 2 to switch 1. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 6 to counter 5.

[0133] In summary, after the tenth switching cycle, a failure occurred between Counter 4 and Counter 5 when Switch 1 transmitted a detection message to Switch 2, and a failure also occurred between Counter 5 and Counter 4 when Switch 2 transmitted a detection message to Switch 1. Therefore, it can be determined that a bidirectional transmission failure occurred between Counter 4 and Counter 5. Because the link between Counter 4 and Counter 5 is an intermediate link, Switch 1 can determine that the fault is an intermediate link failure.

[0134] Example 5: This device has a one-way fault.

[0135] In one possible implementation, Figure 9 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 7.

[0136] Table 7

[0137]

[0138]

[0139] As shown in Table 7, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 and 3, and a count of 10 between counters 4 and 8. Therefore, it can be determined that a transmission failure occurred between counters 3 and 4 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 3 to counter 4.

[0140] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 is forwarded normally. The detection response sent by switch 1 to switch 2 has a count of 15 between counters 1 and 3, and a count of 10 between counters 4 and 8. Therefore, it can be determined that a transmission failure occurred between counters 3 and 4 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to 15th switching cycles failed to be transmitted from counter 3 to counter 4.

[0141] In summary, after the 10th switching cycle, a failure occurred in the test message transmission from Switch 1 to Switch 2 between Counter 3 and Counter 4, while Switch 2 was transmitting test messages to Switch 1 normally. Therefore, it can be determined that a unidirectional transmission failure occurred between Counter 3 and Counter 4. Because the connection between Counter 3 and Counter 4 belongs to Switch 1, Switch 1 can determine that the fault is local to the device.

[0142] Example 6: This device has a bidirectional fault.

[0143] In one possible implementation, Figure 10As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 8.

[0144] Table 8

[0145]

[0146] As shown in Table 8, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 and 3, and a count of 10 between counters 4 and 8. Therefore, it can be determined that a transmission failure occurred between counters 3 and 4 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 3 to counter 4.

[0147] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 has a count of 15 between counter 8 and counter 4, and a count of 10 between counter 3 and counter 1. Therefore, it can be determined that a transmission failure occurred between counter 4 and counter 3 in the direction of transmission from switch 2 to switch 1, and the detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 4 to counter 3.

[0148] In summary, after the tenth switching cycle, a failure occurred between counters 3 and 4 in the test message transmission from Switch 1 to Switch 2, and a failure also occurred between counters 3 and 4 in the test message transmission from Switch 2 to Switch 1. Therefore, it can be determined that a bidirectional transmission failure occurred between counters 3 and 4. Because the connection between counters 3 and 4 belongs to Switch 1, Switch 1 can determine that the fault is local to the device.

[0149] In the above embodiments, there is only one fault point. In addition, there can be multiple fault points, as shown below:

[0150] Example 7: Multiple points of failure.

[0151] In one possible implementation, Figure 11 As shown in the figure, the first device is switch 1, and the second device is switch 2. Switch 1 sends a detection request to switch 2 every second, and switch 2 responds with a corresponding detection reply to switch 1. After 10 cycles, a fault occurs between switches 1 and 2. After 5 detection cycles, switch 1 obtains the detection message statistics of switches 1 and 2, as shown in Table 8.

[0152] Table 9

[0153]

[0154] As shown in Table 9, after the 10th switching cycle, the detection request sent by switch 1 to switch 2 has a count of 15 between counters 1 and 3, and a count of 10 between counters 4 and 8. Therefore, it can be determined that a transmission failure occurred between counters 3 and 4 in the direction of transmission from switch 1 to switch 2. The detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 3 to counter 4.

[0155] After the 10th switching cycle, the detection request sent by switch 2 to switch 1 has a count of 15 between counter 8 and counter 5, and a count of 10 between counter 4 and counter 1. Therefore, it can be determined that a transmission failure occurred between counter 5 and counter 4 in the transmission direction from switch 2 to switch 1, and the detection requests from the 11th to the 15th switching cycles failed to be transmitted from counter 5 to counter 4.

[0156] In summary, after the 10th switching cycle, a failure occurred between counters 3 and 4 in the detection message transmission from switch 1 to switch 2. A transmission failure also occurred between counters 5 and 4 in the direction from switch 2 to switch 1. This indicates a failure on both the local device and the intermediate link.

[0157] In the embodiment of the present application, in addition to the above-mentioned examples, there may be other fault types, and the first device may determine the location information of the fault based on the detection message statistical information, which is not limited here.

[0158] In summary, in the embodiments of the present application, based on the detection message statistics of the first device and the detection message statistics of the second device, the first device can determine whether the fault point is located on the first device (the first device), the opposite device (the second device), or an intermediate link, thereby having a stronger fault location capability. At the same time, it can also determine whether the fault is a unidirectional fault or a bidirectional fault. More fault information can be determined through the detection message, thereby enhancing the ability to analyze the fault.

[0159] 305. Perform fault processing according to the fault location information.

[0160] The first device performs fault processing according to the fault location information, where the fault processing includes operations such as fault resetting and reporting a fault alarm.

[0161] Specifically, after determining the fault location information, the first device may perform a fault reset or report a fault alarm.

[0162] In a possible implementation, when a local device fails, the first device may perform a fault reset or report a fault alarm.

[0163] Specifically, the first device can first determine whether there is continuous queue back pressure on the device. If there is no continuous queue back pressure, the first device can perform fault recovery actions, such as resetting the network chip; it can also report a fault alarm, and the user decides how to handle it later.

[0164] If a queue continues to experience back pressure, the first device only reports a fault alarm and does not perform fault recovery actions.

[0165] In a possible implementation, for an intermediate link failure, the first device may report a failure alarm.

[0166] In a possible implementation, for a failure of the opposite-end device, the first device may not perform any processing, or may report a failure alarm.

[0167] In the embodiment of the present application, although the first device does not perform fault reset for the intermediate link failure and the opposite device failure, after reporting the alarm, the network management system maintenance personnel can further combine the link transceiver module operation status and the port fault status to make a fault judgment, thereby speeding up the fault handling process.

[0168] The embodiment of the present application provides a first device 1200. In the embodiment of the present application, the first device 1200 can be divided into functional modules according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of a software functional module. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0169] In the case of dividing each functional module into corresponding functional modules, Figure 12 FIG. 1 is a schematic diagram showing a possible structure of the first device 1200 involved in the above embodiment. Figure 12 As shown, the first device 1200 includes:

[0170] The sending module 1201 is configured to send a first detection request to the second device, where the first detection request carries detection message statistical information of the first device; for example, step 301, sending the first detection request.

[0171] In a possible implementation, the sending module 1201 is specifically configured to send a first detection request to the second device via the first link. For example, step 301: sending the first detection request.

[0172] The receiving module 1202 is configured to receive a first detection response from the second device, where the first detection response carries detection message statistics information of the second device. For example, in step 302, the first detection response is received.

[0173] In a possible implementation, the receiving module 1202 is specifically configured to: receive a first detection response from the second device through the first link, for example, step 302, receiving the first detection response.

[0174] The acquisition module 1203 is used to obtain the detection message statistical information of the first device and the detection message statistical information of the second device, where the detection message statistical information of the first device is the statistical information of the detection message counter of the first device, and the detection message statistical information of the second device is the statistical information of the detection message counter of the second device, and the detection message is a detection request or a detection response; for example, step 303, obtain the detection message statistical information of the first device and the detection message statistical information of the second device.

[0175] In one possible implementation, the acquisition module 1203 is specifically configured to acquire, via the second link, first device detection message statistics and second device detection message statistics, where the first link and the second link are included in the plurality of communication links. For example, in step 303, the first device detection message statistics and the second device detection message statistics are acquired.

[0176] In a possible implementation, the obtaining module 1203 includes:

[0177] The first acquiring unit 1204 is configured to acquire the first device detection message statistical information through the second detection request; for example, in step 303 , acquire the first device detection message statistical information and the second device detection message statistical information.

[0178] The second acquiring unit 1205 is configured to acquire the second device detection message statistics information through the second detection response, where the second detection request and the second detection response are detection messages corresponding to the second link. For example, in step 303, the first device detection message statistics information and the second device detection message statistics information are acquired.

[0179] The first determining module 1206 is configured to determine fault location information based on the first device detection message statistics and the second device detection message statistics, wherein the fault location information includes: local device failure, opposite device failure, or intermediate link failure. For example, step 304 determines the fault location information.

[0180] In a possible implementation, the first determining module 1206 includes:

[0181] The first determining unit 1207 is configured to determine a transmission interruption point of the detection message according to the first device detection message statistical information and the second device detection message statistical information; for example, in step 304, determine the fault location information.

[0182] The second determining unit 1208 is configured to determine the fault location information according to the transmission interruption point. For example, in step 304, the fault location information is determined.

[0183] The second determining module 1209 is configured to determine whether the fault is a unidirectional fault or a bidirectional fault.

[0184] The execution module 1210 is configured to execute fault processing according to the fault location information. For example, in step 305, the fault processing is executed according to the fault location information.

[0185] In one possible implementation, the execution module 1210 is specifically configured to: if the fault is a peer device fault, report an alarm or not perform any operation; if the fault is an intermediate link fault, report an alarm; or if the fault is a local device fault, report an alarm or perform a fault reset operation. For example, in step 305, fault handling is performed based on the fault location information.

[0186] The various modules of the above-mentioned first device can also be used to perform other actions in the above-mentioned method embodiment. All relevant contents of each step involved in the above-mentioned method embodiment can be referred to the functional description of the corresponding functional module and will not be repeated here.

[0187] The embodiment of the present application provides a second device 1300. In the embodiment of the present application, the second device 1300 can be divided into functional modules according to the above method example. For example, each functional module can be divided corresponding to each function, or two or more functions can be integrated into one processing module. The above-mentioned integrated module can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present invention is schematic and is only a logical functional division. There may be other division methods in actual implementation.

[0188] In the case of dividing each functional module into corresponding functional modules, Figure 13 FIG. 1 is a schematic diagram showing a possible structure of the second device 1300 involved in the above embodiment. Figure 13 As shown, the second device 1300 includes:

[0189] The receiving module 1301 is configured to receive a first detection request from a first device, where the detection request carries detection message statistical information of the first device; for example, in step 301 , the first detection request is sent.

[0190] The sending module 1302 is configured to send a first detection response to the first device, where the detection response carries the detection message statistics information of the second device. For example, in step 302, the first detection response is received.

[0191] The modules of the second device may also be used to perform other actions in the method embodiment. All relevant contents of the steps involved in the method embodiment may be referred to the functional description of the corresponding functional modules and will not be repeated here.

[0192] Figure 14 14 is a schematic diagram of a first device structure provided in an embodiment of the present application. The first device 1400 may include one or more central processing units (CPU) 1401 and a memory 1405. The memory 1405 stores one or more applications or data.

[0193] Memory 1405 may be volatile storage or persistent storage. The program stored in memory 1405 may include one or more modules, each of which may include a series of instruction operations on the first device. Furthermore, the central processing unit 1401 may be configured to communicate with memory 1405 and execute the series of instruction operations in memory 1405 on the first device 1400.

[0194] The central processing unit 1401 is configured to execute a computer program in the memory 1405, so that the first device 1400 is configured to perform the following steps: the first device obtains first device detection message statistical information and second device detection message statistical information, where the first device detection message statistical information is statistical information of the detection message counter of the first device, the second device detection message statistical information is statistical information of the detection message counter of the second device, and the detection message is a detection request or a detection response; and when a transmission failure occurs between the first device and the second device, the first device determines fault location information based on the first device detection message statistical information and the second device detection message statistical information, where the fault location information includes:

[0195] The local device is faulty, the remote device is faulty, or the intermediate link is faulty.

[0196] The first device 1400 may also include one or more power supplies 1402, one or more wired or wireless network interfaces 1403, one or more input and output interfaces 1404, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0197] The first device 1400 can perform the operations performed by the first device in the aforementioned embodiment, and the details are not repeated here.

[0198] Figure 15 15 is a schematic diagram of a second device structure provided in an embodiment of the present application. The second device 1500 may include one or more central processing units (CPU) 1501 and a memory 1505. The memory 1505 stores one or more applications or data.

[0199] Memory 1505 may be volatile or persistent storage. The program stored in memory 1505 may include one or more modules, each of which may include a series of instruction operations on the second device. Furthermore, central processing unit 1501 may be configured to communicate with memory 1505 and execute the series of instruction operations in memory 1505 on the second device 1500.

[0200] Among them, the central processing unit 1501 is used to execute the computer program in the memory 1505, so that the second device 1500 is used to execute: the second device receives a first detection request from the first device, and the first detection request carries the detection message statistical information of the first device; the second device sends a first detection response to the first device, and the first detection response carries the detection message statistical information of the second device.

[0201] The second device 1500 may also include one or more power supplies 1502, one or more wired or wireless network interfaces 1503, one or more input and output interfaces 1504, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.

[0202] The second device 1500 can execute the operations executed by the second device in the aforementioned embodiment, and the details are not repeated here.

[0203] An embodiment of the present application provides a communication system 1600, which includes a first device 1601 and a second device 1602, wherein the first device 1601 can perform the operations performed by the first device in the aforementioned embodiment, and the second device 1602 can perform the operations performed by the second device in the aforementioned embodiment, and the details are not repeated here.

[0204] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes (or functions) of the embodiments of the present application are implemented. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more media that can be integrated. Available media may be magnetic media (eg, floppy disk, hard disk, tape), optical media (eg, DVD), or semiconductor media (eg, solid state disk (SSD)), etc. In the embodiment of the present application, the computer may include the aforementioned device.

[0205] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.

Claims

1. A fault detection method, characterized in that: The method comprises: The first device obtains detection message statistical information of the first device and detection message statistical information of the second device, where the first device detection message statistical information is statistical information of the detection message counter of the first device, and the second device detection message statistical information is statistical information of the detection message counter of the second device, and the detection message is a detection request or a detection response; The first device determines fault location information according to the first device detection message statistical information and the second device detection message statistical information, where the fault location information includes: local device failure, opposite device failure, or intermediate link failure.

2. The method according to claim 1, characterized in that Before the first device obtains the first device detection message statistical information and the second device detection message statistical information, the method further includes: The first device sends a first detection request to the second device, where the first detection request carries detection message statistical information of the first device; The first device receives a first detection response from the second device, where the first detection response carries detection message statistics information of the second device.

3. The method according to claim 2, characterized in that There are multiple communication links between the first device and the second device, and the first device sending a detection request to the second device includes: the first device sending a first detection request to the second device through the first link; The first device receiving a detection response from the second device includes: the first device receiving a first detection response from the second device through the first link; The first device acquiring the first device detection message statistical information and the second device detection message statistical information includes: the first device acquiring the first device detection message statistical information and the second device detection message statistical information through a second link, and the first link and the second link are included in the multiple communication links.

4. The method according to claim 3, characterized in that The detection message includes a detection identifier, a type identifier, a message source address, and a message destination address, wherein the detection identifier is used to indicate that the detection message is a detection message, and the type identifier is used to indicate that the detection message is a detection request or a detection response.

5. The method according to claim 4, characterized in that The first device acquiring the first device detection message statistical information and the second device detection message statistical information includes: The first device obtains first device detection message statistical information through the second detection request; The first device obtains the second device detection message statistical information through the second detection response, where the second detection request and the second detection response are detection messages corresponding to the second link.

6. The method according to claim 5, characterized in that The first device determines fault location information according to the first device detection message statistical information and the second device detection message statistical information, including: The first device determines, according to the statistical information of the detection messages of the first device and the statistical information of the detection messages of the second device, a transmission interruption point of the detection message; The first device determines fault location information according to the transmission interruption point.

7. The method according to claim 6, characterized in that The method further comprises: The first device determines whether the fault is a unidirectional fault or a bidirectional fault.

8. The method according to claim 7, characterized in that The method further comprises: The first device performs fault processing according to the fault location information.

9. The method according to claim 8, characterized in that The first device performs fault processing according to the fault location information, including: The fault location information indicates a fault on the opposite device, and the first device reports an alarm or does not perform an operation; or The fault location information is an intermediate link fault, and the first device reports an alarm; or The fault location information indicates a local device fault, and the first device reports an alarm or performs a fault reset operation.

10. A fault detection method, characterized in that: The method comprises: The second device receives a first detection request from the first device, where the detection request carries detection message statistical information of the first device; The second device sends a first detection response to the first device, where the detection response carries detection message statistics information of the second device.

11. A first device, characterized in that: The first device includes: an acquisition module, configured to acquire detection message statistical information of a first device and detection message statistical information of a second device, wherein the first device detection message statistical information is statistical information of a detection message counter of the first device, and the second device detection message statistical information is statistical information of a detection message counter of the second device, and the detection message is a detection request or a detection response; The first determining module is configured to determine fault location information based on the first device detection message statistical information and the second device detection message statistical information, where the fault location information includes: local device failure, opposite device failure, or intermediate link failure.

12. The first device according to claim 11, characterized in that The first device further includes: A sending module, configured to send a first detection request to the second device, where the first detection request carries detection message statistical information of the first device; The receiving module is configured to receive a first detection response from a second device, where the first detection response carries detection message statistical information of the second device.

13. The first device according to claim 12, characterized in that There are multiple communication links between the first device and the second device, and the sending module is specifically used to: send a first detection request to the second device through the first link; The receiving module is specifically configured to: receive a first detection response from the second device via the first link; The acquisition module is specifically configured to acquire first device detection message statistical information and second device detection message statistical information via a second link, where the first link and the second link are included in the plurality of communication links.

14. The first device according to claim 13, characterized in that The detection message includes a detection identifier, a type identifier, a message source address, and a message destination address, wherein the detection identifier is used to indicate that the detection message is a detection message, and the type identifier is used to indicate that the detection message is a detection request or a detection response.

15. The first device according to claim 14, characterized in that The acquisition module includes: A first acquiring unit, configured to acquire first device detection message statistical information through a second detection request; The second acquiring unit is configured to acquire second device detection message statistical information through a second detection response, where the second detection request and the second detection response are detection messages corresponding to the second link.

16. The first device according to claim 15, characterized in that The first determining module includes: A first determining unit, configured to determine a transmission interruption point of the detection message according to the first device detection message statistical information and the second device detection message statistical information; The second determining unit is configured to determine fault location information according to the transmission interruption point.

17. The first device according to claim 16, characterized in that The first device further includes: The second determining module is configured to determine whether the fault is a unidirectional fault or a bidirectional fault.

18. The first device according to claim 17, characterized in that The first device further includes: An execution module is used to execute fault processing according to the fault location information.

19. The first device according to claim 18, characterized in that The execution module is specifically used for: If the fault is a fault of the opposite device, reporting an alarm or not performing any operation; or If the fault is an intermediate link fault, reporting an alarm; or If the fault is a local device fault, an alarm is reported or a fault reset operation is performed.

20. A second device, characterized in that: The second device includes: A receiving module, configured to receive a first detection request from a first device, wherein the detection request carries detection message statistical information of the first device; The sending module is configured to send a first detection response to the first device, where the detection response carries detection message statistical information of the second device.

21. A first device, characterized in that: The first device includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the first device to implement the method according to any one of claims 1 to 9.

22. A second device, characterized in that: The second device includes a processor and a memory; the processor is coupled to the memory; the memory is used to store computer instructions, and the computer instructions are loaded and executed by the processor to enable the second device to implement the method according to claim 10.

23. A computer-readable storage medium, characterized in that The computer-readable storage medium stores at least one computer program instruction, and the computer program instruction is loaded and executed by a processor to implement the method according to any one of claims 1 to 9.

24. A computer-readable storage medium, characterized in that At least one computer program instruction is stored in the computer-readable storage medium, and the computer program instruction is loaded and executed by the processor to implement the method according to claim 10.

25. A computer program product, characterized in that The computer program product comprises computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer is configured to implement the method according to any one of claims 1 to 9.

26. A computer program product, characterized in that The computer program product comprises computer-executable instructions. When the computer-executable instructions are run on a computer, the computer is configured to implement the method according to claim 10 .

27. A communication system, characterized in that: The communication system includes a first device and a second device, the first device is used to execute the method according to any one of claims 1 to 9, and the second device is used to execute the method according to claim 10.