Fault determination method and device
By counting traffic in the network processor of the communication device to judge port failures, the problem of slow failure determination speed in the prior art is solved, and fast fault determination and efficient path switching are achieved.
Patent Information
- Application Number
- CN202311520353.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-14
- Publication Date
- 2025-05-16
AI Technical Summary
The speed of failure determination in the prior art in communication networks cannot meet the demand, resulting in packet loss of traffic and delayed path switching.
By counting the traffic received by the physical bus in the network processor of the communication device, determining whether the first port is faulty, and rapid failure determination is achieved. This method does not require interaction with other devices to provide protocol messages, saving interaction time and improving fault determination efficiency.
It realizes rapid identification of port failures, reduces service traffic packet loss, improves path switching efficiency, and can complete path switching at microsecond level.
Smart Images

Figure CN120017482A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular to a fault determination method and device. Background Art
[0002] In a communication network, if a link between network devices or a network device fails, the service traffic may not be transmitted along the established path. In order to minimize service traffic packet loss, it is particularly important to quickly detect the above failures so that the network device can quickly switch the service traffic to other paths for forwarding.
[0003] The speed of fault determination of the existing fault determination solutions cannot meet the requirements in some scenarios. Therefore, there is an urgent need for a solution that can solve the above problems. Summary of the invention
[0004] The embodiment of the present application provides a fault determination method, which can quickly determine a port fault.
[0005] In a first aspect, an embodiment of the present application provides a fault determination method, which can be applied to a communication device, wherein the communication device includes a network processor (NP). The communication device also includes a physical bus and a first port, and the physical bus is used to connect the first port and the network processor. The communication device can receive traffic through the first port, and the traffic received by the communication device through the first port can be transmitted to the aforementioned network processor through the physical bus. In an embodiment of the present application, considering the traffic received by the network device through the first port, it can be reflected whether the first port is faulty. For example, under normal circumstances, the traffic received by the network device through the first port is relatively large. When the first port fails, the first port cannot be used to receive traffic, resulting in less traffic received by the network device through the first port. Therefore, in an embodiment of the present application, the network processor can count the traffic received on the physical bus, and the traffic received on the physical bus is the traffic received through the first port. Specifically, the network processor can count the traffic received on the physical bus before parsing the traffic received on the physical bus to obtain a statistical result, and determine the first port fault according to the statistical result. It can be seen that, by using the solution of the embodiment of the present application, the network device can determine the fault of the first port without exchanging protocol messages with other devices, thereby saving the time of interacting with other devices and correspondingly improving the efficiency of determining the fault of the first port. In addition, the network processor counts the traffic before parsing the traffic, which also saves the time required for parsing the traffic. Therefore, by using this solution, the fault of the first port can be quickly determined.
[0006] In a possible implementation, when determining whether the first port is faulty, the network processor may determine the faulty port by combining the statistical results corresponding to each statistical period in a plurality of consecutive statistical periods. Specifically, if the traffic counted in each statistical period in the aforementioned plurality of consecutive statistical periods is less than the first traffic threshold, it means that the traffic received by the first port is very small during the period corresponding to the aforementioned plurality of consecutive statistical periods. In this case, it can be considered that the first port is faulty.
[0007] In a possible implementation, the network processor may perform statistics on the traffic received on the physical bus when the network device has enabled the rapid fault detection function. As a specific example, the network processor may run a state machine when the network device has enabled the rapid fault detection function, and the state machine may include multiple states, each state having a corresponding execution action. Among them, the multiple states include a fault detection state, and the execution action corresponding to the fault detection state may include counting the traffic and determining whether the first port is faulty based on the statistical results obtained by the statistics. In other words, the network processor may count the traffic in the fault detection state of the state machine, and further, in the fault detection state, determine whether the first port is faulty based on the statistical results obtained by the statistics.
[0008] In one possible implementation, after the network processor determines that the first port is faulty in the fault detection state of the state machine, the state of the state machine is switched from the fault detection state to the fault confirmation state, so that the network processor can further execute the execution action of the fault confirmation state, and no longer continue to execute the action of "counting traffic and determining whether the first port is faulty based on the statistical results obtained by statistics".
[0009] In a possible implementation, considering that under normal circumstances, the network device receives a relatively large amount of traffic through the first port, when the first port fails, the first port cannot be used to receive traffic, resulting in the network device receiving a relatively large amount of traffic through the first port. Therefore, in one example, before counting the traffic received by the physical bus, the network processor may also determine that the traffic received by the first port within a preset time period is greater than the second traffic threshold, thereby determining that the traffic received by the first port has undergone a change process of "from more to less" or "from existence to nothing". Specifically, before counting the traffic received by the physical bus, the network processor may determine that the traffic received by the physical bus within a preset time period is greater than the second traffic threshold.
[0010] In a possible implementation, the network processor can count the traffic received by the physical bus in the initial state of the state machine, and determine that the traffic received by the first port within a preset time period is greater than the second traffic threshold. That is, in the initial state of the state machine, the network processor can determine that the first port "has" traffic (or determines that the first port receives more traffic). In one example, after determining that the traffic received by the first port within a preset time period is greater than the second traffic threshold, the state of the state machine can be switched from the initial state to the aforementioned fault detection state, so as to continue to determine the change trend of the traffic received by the first port in the aforementioned fault detection state, and determine that the first port is faulty when the change trend of the traffic received by the first port is "from more to less" or "from there to nothing".
[0011] In a possible implementation, after the network processor determines that the first port is faulty, it can save fault indication information indicating the fault of the first port, so as to perform path switching based on the fault indication information, thereby achieving fast path switching. In one example, using the solution of the embodiment of the present application, path switching at the microsecond level can be achieved, effectively reducing packet loss.
[0012] In a possible implementation, the network processor may receive a first service message. After receiving the first service message, the network processor may search the first forwarding table. According to the first forwarding table, the network processor may determine that the outgoing port for forwarding the first service message includes the first port and the second port. After determining the first port and the second port, the network processor may select to forward the first service message through the second port according to the aforementioned fault indication information. That is, in the event of a fault in the first port, the second port is selected to forward the first service message, thereby avoiding the first service message being lost due to forwarding the first service message through the first port.
[0013] In a possible implementation, considering that in actual applications, the first port may also have a short-term interruption phenomenon, when the first port has a short-term interruption, the traffic received by the first port in one or more continuous statistical cycles will also be less than the first traffic threshold. However, in this case, it is unreasonable to determine that the first port is faulty, that is, determining that the first port fault may be a false detection. In order to solve this problem, the network processor can determine whether a false detection has occurred in combination with the result of the central processing unit (CPU) detecting the state of the first port. In an example, the network processor can clear the fault indication information when the storage time of the aforementioned fault indication information reaches a preset time. Among them, the preset time is greater than the time for the central processor of the communication device to determine the port state of the first port. Among them: after the central processor detects the state of the first port, it can be combined with the state of the first port to select whether to update the first forwarding table. Specifically, if the first port detects that the state of the first port is up, the central processor may not update the first forwarding table, and accordingly, the first forwarding table also includes a forwarding table entry whose outgoing port is the first port. At this time, since the aforementioned fault indication information is cleared, the network processor can subsequently continue to use the first port to forward the service message, thereby realizing the rapid switching back of the path. If the first port detects that the state of the first port is closed (down), the central processor can update the first forwarding table. Specifically, the forwarding table entry corresponding to the first port in the first forwarding table can be deleted to obtain the second forwarding table entry. Therefore, the fault indication information is cleared, and the service message will not be forwarded through the faulty first port. In this way, whether the first port fault is determined to be a false detection or not, it will not affect the forwarding of the service traffic.
[0014] In a possible implementation, the network processor may save fault indication information indicating the fault of the first port in the fault confirmation state of the state machine. Furthermore, in the fault confirmation state, when the storage time of the fault indication information reaches the aforementioned preset time, the network processor may clear the aforementioned fault indication information. In addition, when the storage time of the fault indication information reaches the aforementioned preset time, the state of the state machine is switched from the fault confirmation state to the initial state of the state machine, so as to start a new round of rapid fault detection process.
[0015] In a possible implementation, if the central processor determines that the state of the first port is down, it means that the first port fault is not a false detection. In this case, the central processor deletes the forwarding table entry corresponding to the first port in the first forwarding table entry to obtain the second forwarding table entry. Therefore, the fault indication information is cleared, and the service message will not be forwarded through the first port of the fault. As a specific example, the network processor can receive the second service message. After the network processor receives the second service message, it can search the second forwarding table. In a specific example, the second forwarding table includes a forwarding table entry with the second service message, and the outgoing port of the forwarding table entry is the second port. After determining the second port, the network processor can forward the second service message through the second port. That is: when it is determined that the first port fault is not a false detection, although the fault indication information is cleared, since the central processor updates the first forwarding table, the second service message can actually be forwarded through the second port without fault.
[0016] In a possible implementation, if the central processor determines that the state of the first port is up, it indicates that the determination of the first port fault is a false detection. In this case, since the fault indication information is cleared, the network processor can continue to use the first port to forward service messages, thereby realizing rapid path switching. As a specific example, the network processor can receive a third service message. After receiving the third service message, the network processor can search the first forwarding table to determine that the outgoing port for forwarding the third service message includes the first port and the second port. After determining the first port and the second port, the network processor can forward the second service message through the first port.
[0017] In a possible implementation, the communication device is a box-type network device, which can be considered as a device including a service board.
[0018] In a possible implementation manner, the communication device is a first service board in a frame-type network device.
[0019] In a possible implementation, if the communication device is a first service board in a frame network device, after determining that the first port fails, the first service board can further notify the first port failure to at least one second service board of the frame network device through the network board of the frame network device, so that the second service board can perform path switching based on the first port failure, thereby enabling the entire frame network device to achieve fast path switching. In a specific example, the first service board can send a broadcast message to at least one second service board through the network board, and the broadcast message indicates the failure of the first port.
[0020] In a second aspect, an embodiment of the present application provides a network processor, which belongs to a communication device, and the network processor includes: a processing unit; the processing unit is used to count the traffic received on the physical bus to obtain statistical results before parsing the traffic, and the communication device includes the physical bus and a first port, and the physical bus is used to connect the first port and the network processor; based on the statistical results, it is determined that the first port is faulty.
[0021] In a possible implementation, the statistical result includes a statistical result corresponding to each statistical period in a plurality of consecutive statistical periods, and the processing unit is used to determine that the first port is faulty when the traffic counted in each statistical period is less than a first traffic threshold.
[0022] In a possible implementation, the processing unit is used to: run a state machine when a rapid fault detection function is enabled; count the traffic in a fault detection state of the state machine; and determine the first port fault based on the statistical result in the fault detection state.
[0023] In a possible implementation manner, after determining that the first port is faulty, the state of the state machine is switched from the fault detection state to the fault confirmation state.
[0024] In a possible implementation, the processing unit is further configured to: before counting the traffic, determine whether the traffic received by the physical bus within a preset time period is greater than a second traffic threshold.
[0025] In one possible implementation, the determining that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold includes: in the initial state of the state machine, determining that the traffic received by the physical bus within a preset time period is greater than the second traffic threshold; wherein: after determining in the initial state that the traffic received by the physical bus within the preset time period is greater than the second traffic threshold, the state of the state machine is switched from the initial state to the fault detection state.
[0026] In a possible implementation manner, the processing unit is further configured to: save fault indication information indicating a fault of the first port.
[0027] In one possible implementation, the network processor also includes: a receiving unit and a sending unit; the receiving unit is used to receive a first service message; the processing unit is also used to determine, based on a first forwarding table, that an egress port for forwarding the first service message includes the first port and a second port, and the communication device includes the second port; the sending unit is used to select, based on the fault indication information, to forward the first service message through the second port.
[0028] In a possible implementation, the processing unit is further used to: clear the fault indication information when the storage time of the fault indication information reaches a preset time, and the preset time is greater than the time for the central processor of the communication device to determine the port status of the first port.
[0029] In one possible implementation, after determining that the first port fault occurs, when the state of the state machine is switched from the fault detection state of the state machine to the fault confirmation state, the processing unit saves the fault indication information in the fault confirmation state, and when the storage time of the fault indication information reaches the preset time, the state of the state machine is switched from the fault confirmation state to the initial state of the state machine.
[0030] In one possible implementation, after clearing the fault indication information, if the central processor determines that the status of the first port is closed, the network processor includes a receiving unit that is also used to: receive a second service message; the processing unit is also used to determine, based on a second forwarding table, that the egress port for forwarding the second service message includes the second port, wherein the second forwarding table is a forwarding table obtained by the central processor by updating the first forwarding table when determining that the status of the first port is closed; the network processor includes a sending unit that is also used to forward the second service message through the second port.
[0031] In one possible implementation, after clearing the fault indication information, if the central processor determines that the status of the first port is on, the network processor includes a receiving unit that is also used to: receive a third business message; the processing unit is also used to determine, based on the first forwarding table, that the egress port for forwarding the third business message includes the first port and the second port; the network processor includes a sending unit that is also used to select to forward the second business message through the first port.
[0032] In a possible implementation manner, the communication device is a box-type network device.
[0033] In a possible implementation manner, the communication device is a first service board in a frame-type network device.
[0034] In one possible implementation, the frame network device also includes a second business board and a network board, the network board is connected to each business board in the frame network device, and the each business board includes the first business board and the second business board. The network processor includes a sending unit, which is also used to: send a broadcast message to the second business board through the network board, and the broadcast message indicates a failure of the first port so that the second business board can perform path switching based on the broadcast message.
[0035] In a third aspect, an embodiment of the present application provides a device comprising: a processor and a memory; the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs, and execute the method described in the first aspect and any one of the above first aspects.
[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, including instructions or a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect and any one of the above first aspects. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0038] Figure 1a is a schematic diagram of another exemplary application scenario;
[0039] Figure 1b is a schematic diagram of another exemplary application scenario;
[0040] Figure 1c is a schematic diagram of another exemplary application scenario;
[0041] Figure 2a A schematic diagram of the structure of a box-type network device provided in an embodiment of the present application;
[0042] Figure 2b A schematic diagram of the structure of a frame-type network device provided in an embodiment of the present application;
[0043] Figure 3 A flowchart of a fault determination method provided in an embodiment of the present application;
[0044] Figure 4 A schematic diagram of a state machine provided in an embodiment of the present application;
[0045] Figure 5 A schematic diagram of a process of rapid fault detection and path switching provided in an embodiment of the present application;
[0046] Figure 6 A schematic diagram of another process of rapid fault detection and path switching provided in an embodiment of the present application;
[0047] Figure 7 A schematic diagram of the structure of a network processor provided in an embodiment of the present application;
[0048] Figure 8 A schematic diagram of the structure of a device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The embodiment of the present application provides a fault determination method, which can quickly determine a port fault.
[0050] At present, fault detection can be performed by deploying a bidirectional forwarding detection (BFD) protocol in the network, and when a fault is detected, the service traffic is switched to other paths for forwarding, thereby avoiding service traffic packet loss. Using the BFD protocol for fault detection requires the exchange of BFD protocol messages between communication devices. The BFD protocol messages mentioned here may include the BFD detection messages and BFD response messages mentioned below. In some scenarios, the BFD detection message may also be referred to as a BFD request message.
[0051] Next, the principle of using the BFD protocol for fault detection is introduced by exchanging BFD protocol packets between communication device A and communication device B.
[0052] Communication device A periodically sends a BFD detection message to communication device B, and determines whether a fault occurs according to whether a BFD response message returned by communication device B is received. As an example, if communication device A receives a BFD response message returned by communication device B within a certain period of time, it is determined that the link between communication device A and communication device B has not failed. If communication device A does not receive a BFD response message returned by communication device B within a certain period of time, it is determined that a fault occurs in the link between communication device A and communication device B.
[0053] In some scenarios, considering that in the case of congestion on the link between communication device A and communication device B, communication device A may not be able to receive the BFD response message returned by communication device B within the aforementioned certain time period, therefore, for a certain BFD detection message sent by communication device A to communication device B, if communication device A does not receive the BFD response message returned by communication device B for the BFD detection message within a certain time period, it does not mean that the link between communication device A and communication device B is really faulty. In order to improve the accuracy of fault detection, in a specific example, if communication device A does not receive the BFD response message returned by communication device B for each of the multiple BFD detection messages continuously sent by communication device A to communication device B within a certain time period, it is determined that the link between communication device A and communication device B is faulty.
[0054] As described above, it can be seen that the use of the BFD protocol for fault detection requires the exchange of BFD protocol messages between network devices, and the exchange of BFD protocol messages between communication devices will also occupy some network resources. In order to prevent BFD protocol messages from occupying too many network resources, the period of the aforementioned communication device A periodically sending BFD detection messages to communication device B will not be set very small. Generally, the period can be set to about 10 milliseconds (ms). Therefore, the detection time for fault detection using the BFD protocol is about tens of milliseconds. For example:
[0055] Assuming that for each of the three BFD detection messages sent continuously from communication device A to communication device B, communication device A has not received a BFD response message returned by communication device B for the BFD detection message within a certain period of time, it is determined that a failure has occurred in the link between communication device A and communication device B. It takes about 20 to 30 milliseconds for communication device A to determine that a failure has occurred in the link between communication device A and communication device B. Correspondingly, the switching time from determining the link failure between communication device A and communication device B to switching the service traffic to other paths for forwarding also takes tens of milliseconds. In other words, the BFD protocol can achieve path switching at the level of tens of milliseconds. During the path switching process, packet loss will occur in the service traffic that was originally supposed to be transmitted between network device A and network device B.
[0056] However, in some scenarios, for example, for services that are sensitive to packet loss, path switching at the level of tens of milliseconds cannot meet service requirements.
[0057] Therefore, there is an urgent need for a solution that can quickly detect faults so as to quickly switch paths.
[0058] The inventor of the present application has found that since it takes a certain amount of time to exchange protocol messages between communication devices, if fault detection can be implemented without exchanging protocol messages between communication devices, the efficiency of fault detection can be greatly improved.
[0059] In one scenario, a link failure can be determined by detecting the state of a port. The state of a port may include up and down. As a specific example, the state of a port may be detected by a CPU of a communication device. It may take several milliseconds for the CPU to detect the state of a port.
[0060] With this solution, it is possible to achieve millisecond-level path switching in the data center network. This is because the transmission distance between switches in the data center network is short. Figure 1a To understand, Figure 1a FIG. 1 is a schematic diagram of an exemplary application scenario. Figure 1a As shown, port 1 / 0 / 0 of switch S1 and port 2 / 0 / 0 of switch S2 are directly connected. In this case, if port 1 / 0 / 0 of switch S1 fails, switch S1 can trigger a fault notification in the data center network after detecting that the status of its own port 1 / 0 / 0 is down, so that the switches in the data center network can update the forwarding table, so that the switches can forward based on the updated forwarding table. Since the transmission distance between switches in the data center network is short, the above-mentioned fault notification is also short in time. Therefore, in this way, fast path switching can be achieved in the data center network.
[0061] However, in network scenarios where the transmission distance between communication devices is relatively long (such as operator networks), millisecond-level path switching cannot be achieved. This is because if there are other transmission devices between the communication devices, the aforementioned fault notification takes a long time, and accordingly, the fault switching time is also relatively long. Figure 1b The application scenario shown is described below.
[0062] Figure 1b This is a schematic diagram of another exemplary application scenario.
[0063] like Figure 1bAs shown, a transmission device T1 and a transmission device T2 are included between router R1 and router R2. Router R1 includes port 1 / 0 / 0, and router R2 includes port 2 / 0 / 0. Router R1 can communicate with transmission device T1 through its own port 1 / 0 / 0, and router R2 can communicate with transmission device T2 through its own port 2 / 0 / 0. In one example, transmission devices T1 and T2 can be optical transport network (OTN) devices, for example, router R1 and router R2 are both backbone network devices, and transmission device T1 and transmission device T2 can be OTN devices. In another example, transmission devices T1 and T2 can be switches or routers, for example, router R1 and router R2 are both access devices, and transmission device T1 and transmission device T2 can be switches or routers.
[0064] In one example, if port 1 / 0 / 0 of router R1 fails, router R1 can send a remote fault notification message to router R2 after detecting that the status of its own port 1 / 0 / 0 is down, so that router R2 can determine the occurrence of a remote fault according to the remote fault notification message. However, since transmission equipment T1 and transmission equipment T2 are included between router R1 and router R2, the notification time of the remote fault notification message will be relatively long. Alternatively, router R2 can also determine the link failure between router R1 and router R2 through the aforementioned BFD protocol. However, as can be seen from the above description of fault detection using the BFD protocol, it will take a relatively long time for router R2 to determine the link failure between router R1 and router R2 through the BFD protocol.
[0065] In view of this, the embodiment of the present application provides a fault determination method applicable to various networks, which can quickly detect faults so as to quickly switch paths. The various networks mentioned here include but are not limited to operator networks and data center networks.
[0066] In a specific example, the method provided in the embodiment of the present application can be applied to Figure 1a or Figure 1b The application scenario shown.
[0067] In another specific example, the method provided in the embodiment of the present application can also be applied to Figure 1c The application scenario shown, Figure 1c FIG. 1 is a schematic diagram of another exemplary application scenario. Figure 1c In the scenario shown, port 1 / 0 / 0 of router R1 and port 2 / 0 / 0 of router R2 are directly connected.
[0068] The fault determination method provided in the embodiment of the present application can be applied to a communication device, which can be a switch or a router. Figure 1a In the application scenario shown, the communication device may be a switch S1 or S2. Figure 1b In the application scenario shown, the communication device may be a router R1 or R2. When the method is applied to Figure 1c In the application scenario shown, the communication device may be a router R1 or R2.
[0069] In the embodiment of the present application, the communication device may be a box-type network device or a service board in a frame-type network device. A box-type network device may be considered as a device including a service board, while a frame-type network device is a device including multiple service boards. In addition to multiple service boards, a frame-type network device also includes a network board, and multiple service boards of a frame-type network device may interact with each other through the network board.
[0070] Before introducing the fault determination method provided in the embodiment of the present application, first Figure 2a and Figure 2b This section introduces the structure of box-type network devices and frame-type network devices.
[0071] See also Figure 2a , Figure 2a A schematic diagram of the structure of a box-type network device provided in an embodiment of the present application.
[0072] like Figure 2a As shown, the box-type network device includes: a port chip 210, a network processor 220, a traffic management (TM) chip 230 and a central processor 240. Among them:
[0073] The port chip 210 includes a port, which can be used to receive and / or send data. In one example, the port chip can be a physical layer (PHY) chip or a MAC chip, which is not specifically limited in the embodiments of the present application. The PHY chip is used to implement physical layer related functions such as optical and electrical signals required for data transmission; the MAC chip is used to implement link layer related functions such as data frame construction, data error checking and transmission control. In another example, the port chip 210 includes a card slot that can be inserted into an optical module.
[0074] The port chip 210 and the network processor 220 may be connected via a physical bus.
[0075] The network processor 220 is used to perform forwarding plane related functions, wherein the forwarding plane related functions include but are not limited to: message parsing, message routing, executing access control list (ACL) technology, reporting relevant information to the central processor 240, and other functions.
[0076] The traffic management chip 230 is used to perform traffic control related functions such as congestion management and port scheduling.
[0077] The central processor 240 is used to execute the related functions of the control plane, where the related functions of the control plane include but are not limited to: running the control protocol, updating the message forwarding table, etc.
[0078] See also Figure 2b , Figure 2b A schematic diagram of the structure of a frame-type network device provided in an embodiment of the present application.
[0079] A frame-type network device may include multiple service boards ( Figure 2b The diagram shows two service boards) and a network board, each of which may include a port chip, a network processor, a traffic management chip, and a central processing unit. The traffic management chip of each service board may be connected to the network board, thereby enabling multiple service boards to interact with each other through the network board. For the port chip, the network processor, the traffic management chip, and the central processing unit, please refer to the above description of Figure 2a The description part will not be repeated here.
[0080] Next, combine Figure 3 , the fault determination method provided in the embodiment of the present application is introduced. Figure 3 , which is a flow chart of a fault determination method provided in an embodiment of the present application. Figure 3 The method shown may include the following S101-S102.
[0081] S101: Before parsing the traffic received on the physical bus, the network processor of the communication device counts the traffic to obtain statistical results, wherein the communication device includes the physical bus and a first port, and the physical bus is used to connect the first port and the network processor.
[0082] In an embodiment of the present application, the communication device includes a port chip, and the port chip includes the first port. Regarding the port chip, reference can be made to the relevant description above, and no repeated description is made here. In an embodiment of the present application, in addition to the port chip and the network processor, the communication device also includes a physical bus. The physical bus is used to connect the first port and the network processor. In other words, the data exchanged between the first port and the network processor can be transmitted through the physical bus.
[0083] In an embodiment of the present application, the communication device can receive traffic through the first port, and the traffic received by the communication device through the first port can be transmitted to the network processor via the physical bus, and the network processor counts the traffic. In an embodiment of the present application, considering the traffic received by the network device through the first port, whether the first port is faulty can be reflected. For example, under normal circumstances, the network device receives a lot of traffic through the first port. When the first port fails, the first port cannot be used to receive traffic, resulting in less traffic received by the network device through the first port. Therefore, in an embodiment of the present application, the network processor can count the traffic received by the first port. As mentioned above, the traffic received by the communication device through the first port can be transmitted to the network processor via the physical bus. Therefore, in a specific example, the network processor can count the traffic received on the physical bus, and determine whether the first port is faulty based on the statistical results obtained by the statistics. In one example, in order to maximize the efficiency of fault determination, the network processor can count the traffic received on the physical bus before parsing the traffic.
[0084] In one example, the network processor may use a certain duration as a statistical period, and count the traffic received on the physical bus in each statistical period. For example, 300 microseconds (us) may be used as a statistical period, and the traffic received on the physical bus may be counted in each statistical period. In an embodiment of the present application, counting the traffic received on the physical bus may be counting the number of messages received on the physical bus, or may be counting the size of messages received on the physical bus, which is not specifically limited in the embodiment of the present application.
[0085] In one example, the network processor may execute S101 when the network device has enabled the rapid fault detection function. The network device may enable the rapid fault detection function by default, or may be configured to enable the rapid fault detection function by configuration, which is not specifically limited in the present embodiment.
[0086] As a specific example, the network processor may run a state machine when the network device has enabled a rapid fault detection function, and the state machine may include multiple states, each of which has a corresponding execution action. The multiple states include a fault detection state, and the execution action corresponding to the fault detection state may include counting traffic and determining whether the first port is faulty based on the statistical results obtained by counting. In other words, the network processor may count the traffic in the fault detection state of the state machine, and further, in the fault detection state, determine whether the first port is faulty based on the statistical results obtained by counting. As a specific example, as long as the state machine is in the fault detection state, the network processor periodically executes the action of "counting traffic and determining whether the first port is faulty based on the statistical results obtained by counting".
[0087] S102: The network processor of the communication device determines that the first port is faulty according to the statistical result.
[0088] As mentioned above, after the network processor counts the traffic, it can further determine whether the first port is faulty according to the statistical results. In a specific example, the network processor can determine that the first port is faulty according to the statistical results. The network processor can use a certain time length as a statistical period, and count the traffic received on the physical bus in each statistical period. In an example, the network processor can determine that the first port is faulty when the traffic counted in the current statistical period is less than the first traffic threshold. The first traffic threshold mentioned here can be, for example, a message number threshold. For example, the first port fault can be determined when the number of messages counted in the current statistical period is less than 1. The first traffic threshold mentioned here can also be a message size threshold. For example, the first port fault can be determined when the size of the message counted in the current statistical period is less than 1 megabyte. In another example, considering that the traffic counted in a statistical period is not representative, the network processor can determine whether the first port is faulty by combining the statistical results corresponding to each statistical period in multiple consecutive statistical periods. Specifically, if the traffic counted in each of the aforementioned multiple continuous statistical cycles is less than the first traffic threshold, it means that the traffic received by the first port is very small during the period corresponding to the aforementioned multiple continuous statistical cycles. In this case, it can be considered that the first port has a fault. In other words, the network processor can determine that the first port has a fault when the traffic counted in each of the aforementioned multiple continuous statistical cycles is less than the first traffic threshold. In one example, if the network processor executes S101-S102 in the fault detection state of the state machine. After the network processor determines that the first port has a fault, the state of the state machine is switched from the fault detection state to the fault confirmation state, so that the network processor no longer continues to execute the action of "counting the traffic and determining whether the first port is faulty based on the statistical results obtained from the statistics".
[0089] As mentioned above, under normal circumstances, the network device receives a relatively large amount of traffic through the first port. When the first port fails, the first port cannot be used to receive traffic, resulting in a relatively large amount of traffic received by the network device through the first port. Therefore, in one example, before executing S101, the network processor may also determine that the traffic received by the first port within a preset time period is greater than the second traffic threshold, thereby determining that the traffic received by the first port has undergone a change process of "from more to less" or "from there to nothing". Specifically, the network processor may determine that the traffic received by the physical bus within a preset time period is greater than the second traffic threshold.
[0090] The embodiment of the present application does not specifically limit the preset time period, and the duration corresponding to the preset time period may be, for example, 1 second.
[0091] In one example, the network processor may count the total traffic received by the physical bus within the preset time period, and further determine that the total traffic is greater than the aforementioned second traffic threshold. In another example, the network processor may use a certain time length as a statistical period, and obtain the traffic counted in each statistical period within the preset time period. Further, it is determined that the traffic counted in each statistical period within the preset time period is greater than the second traffic threshold. The second traffic threshold is similar to the first traffic threshold, and it may also be a message quantity threshold or a message size threshold. In a specific example, the second traffic threshold is greater than the aforementioned first traffic threshold.
[0092] As mentioned above, the network processor can run a state machine when the network device has enabled the fast fault detection function. The state machine can include multiple states, and each state has a corresponding execution action. In one example, when the network processor just starts running the state machine, the state machine is in the initial state. The action executed in the initial state may include the aforementioned "determining that the traffic received by the first port within a preset time period is greater than the second traffic threshold". In other words, in the initial state of the state machine, the network processor can count the traffic received by the physical bus and determine that the traffic received by the first port within a preset time period is greater than the second traffic threshold. That is: in the initial state of the state machine, the network processor can determine that the first port "has" traffic (or determines that the first port receives more traffic). In one example, after determining that the traffic received by the first port within a preset time period is greater than the second traffic threshold, the state of the state machine can be switched from the initial state to the aforementioned fault detection state, so as to continue to determine the change trend of the traffic received by the first port in the aforementioned fault detection state, and determine that the first port is faulty when the change trend of the traffic received by the first port is "from more to less" or "from there to nothing".
[0093] It can be seen from the above description that, by using the solution of the embodiment of the present application, the network device can determine the fault of the first port without exchanging protocol messages with other devices, thereby saving the time of interacting with other devices and correspondingly improving the efficiency of determining the fault of the first port. In addition, the network processor counts the traffic before parsing the traffic, which also saves the time required for parsing the traffic. Therefore, by using this solution, the fault of the first port can be quickly determined.
[0094] In one example, after the network processor determines that the first port has failed, it can save fault indication information indicating the first port failure, so as to facilitate path switching based on the fault indication information. In a specific example, the network processor can store the fault indication information in a forwarding table entry corresponding to the first port in the message forwarding table. Among them, the forwarding table entry corresponding to the first port can be a forwarding table entry whose egress port is the first port. For example, a fault indication bit can be added to the forwarding table entry. For a certain forwarding table entry, the fault indication bit corresponding to the forwarding table entry (for example, a value of 1) is set, indicating that the egress port in the forwarding table entry is faulty. This can be understood in conjunction with the following Table 1. Table 1 shows a portion of a forwarding table entry.
[0095] Table 1
[0096] Outgoing port Fault indication bit m1 1 m2 0 m3 0 m4 0
[0097] As shown in Table 1, the value of the fault indication bit of the egress port m1 is 1, indicating that the egress port m1 is faulty, and the values of the fault indication bits of the egress ports m2, m3 and m4 are 0, indicating that the egress ports m2, m3 and m4 are not faulty.
[0098] As other contents not shown in the forwarding table (such as prefix, next hop, etc.) are not closely related to the core idea of the present application, they are not described in detail here.
[0099] As described above, after the network processor determines that the first port is faulty, the state of the state machine is switched from the fault detection state to the fault confirmation state. In one example, the execution action corresponding to the fault determination state may include the aforementioned saving of the fault indication information indicating the fault of the first port. In other words, the network processor may save the fault indication information indicating the fault of the first port in the fault confirmation state of the state machine.
[0100] Next, a specific implementation method of the network processor performing path switching based on the fault indication information is described.
[0101] In one example, the network processor may receive a first service message. After receiving the first service message, the network processor may search a first forwarding table. In a specific example, the first forwarding table includes multiple forwarding table entries related to the first service message, wherein the egress port of one of the forwarding table entries is the first port, and the egress port of another forwarding table entry is the second port. In one example, the first port and the second port may be mutually equivalent load-sharing egress ports. After determining the first port and the second port, the network processor may select to forward the first service message through the second port according to the aforementioned fault indication information. That is, in the event of a fault in the first port, the second port is selected to forward the first service message, thereby avoiding the first service message being lost due to forwarding the first service message through the first port.
[0102] In one example, considering that in actual applications, the first port may also have a short-term interruption phenomenon, when the first port has a short-term interruption, the traffic received by the first port in one or more continuous statistical cycles will also be less than the first traffic threshold. However, in this case, it is unreasonable to determine that the first port is faulty, that is, determining that the first port fault may be a false detection. In order to solve this problem, the network processor can determine whether a false detection has occurred in combination with the result of the central processor detecting the state of the first port. In an embodiment of the present application, it takes a certain time for the central processor to detect the state of the first port. After the central processor detects the state of the first port, it can select whether to update the first forwarding table in combination with the state of the first port. Specifically, if the first port detects that the state of the first port is up, the central processor may not update the first forwarding table, and accordingly, the first forwarding table also includes a forwarding table entry whose outgoing port is the first port. If the first port detects that the state of the first port is down, the central processor can update the first forwarding table, and specifically, the forwarding table entry corresponding to the outgoing port of the first port in the first forwarding table can be deleted to obtain a second forwarding table entry. In one example, the network processor can clear the fault indication information when the storage time of the aforementioned fault indication information reaches a preset time. The preset time duration is greater than the time duration for the central processor of the communication device to determine the port status of the first port.
[0103] So:
[0104] On the one hand, if the central processor determines that the state of the first port is down, it means that the first port fault is not a false detection. In this case, the central processor will delete the forwarding table item corresponding to the first port in the first forwarding table item to obtain the second forwarding table item. Therefore, the fault indication information is cleared, and the service message will not be forwarded through the first port of the fault. As a specific example, the network processor can receive the second service message. After the network processor receives the second service message, it can search the second forwarding table. In a specific example, the second forwarding table includes a forwarding table item with the second service message, and the outgoing port of the forwarding table item is the second port. After determining the second port, the network processor can forward the second service message through the second port. That is: when it is determined that the first port fault is not a false detection, although the fault indication information is cleared, since the central processor updates the first forwarding table, the second service message can actually be forwarded through the second port without fault.
[0105] On the other hand, if the central processor determines that the state of the first port is up, it means that the determination of the first port fault is a false detection. In this case, since the fault indication information is cleared, the network processor can continue to use the first port to forward service messages, thereby realizing rapid path switching. As a specific example, the network processor can receive a third service message. After the network processor receives the third service message, it can search the first forwarding table. In a specific example, the first forwarding table includes multiple forwarding table entries related to the first service message, wherein the egress port of one of the forwarding table entries is the first port, and the egress port of another forwarding table entry is the second port. In an example, the first port and the second port can be equivalent load sharing egress ports to each other. After determining the first port and the second port, the network processor can forward the second service message through the first port.
[0106] As mentioned above, the network processor can save the fault indication information indicating the fault of the first port in the fault confirmation state of the state machine. After the network processor saves the fault indication information in the fault confirmation state, it can further start timing in the fault confirmation state, and when the timing reaches the preset time length, the network processor can clear the aforementioned fault indication information. In addition, the state of the state machine is switched from the fault confirmation state to the initial state of the state machine, so as to start a new round of fault rapid detection process.
[0107] As mentioned above, the communication device may be a box-type network device or a service board in a frame-type network device. In a specific example, if the communication device is a first service board in a frame-type network device, after determining that the first port fails, the first service board may further notify the first port failure to at least one second service board of the frame-type network device through the network board of the frame-type network device, so that the second service board can perform path switching based on the first port failure, thereby enabling the entire frame-type network device to achieve fast path switching. Wherein, when the second service board performs path switching based on the first port failure, for example, the network processor of the second service board sends the service message originally sent to the first port to other ports. In a specific example, the first service board may send a broadcast message to at least one second service board through the network board, and the broadcast message indicates the failure of the first port. In a specific example, the broadcast message may include the identifier of the first port and the indication information indicating the failure of the first port. Wherein, the identifier of the first port may, for example, include the identifier of the first service board to which the first port belongs and the port number of the first port.
[0108] In an example, the network processor may send a broadcast message to the at least one second service board through the network board in the fault confirmation state of the aforementioned state machine.
[0109] Next, combine Figure 4 The state machine shown illustrates the fault confirmation method provided in the embodiment of the present application. Figure 4 A schematic diagram of a state machine provided in an embodiment of the present application.
[0110] like Figure 4 As shown, the state machine includes three states, namely: initial state, fault detection state and fault confirmation state.
[0111] When the network processor determines that the fault fast detection function is enabled, the state machine is run. When the state machine starts running, it is in the initial state.
[0112] The network processor counts the traffic received by the physical bus in the initial state of the state machine. When the network processor determines that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold, the state machine jumps to a fault detection state.
[0113] The network processor, in the fault detection state of the state machine, counts the traffic received on the physical bus before parsing the traffic. If the first port is determined to be faulty according to the statistical result, the state machine jumps to the fault confirmation state.
[0114] The network processor saves the fault indication information of the first port in the fault confirmation state of the state machine. In addition, the storage time of the fault indication information is timed, and once the storage time reaches a preset time, the fault indication information is cleared, and when the storage time of the fault indication information stored by the network processor reaches the preset time, the state machine jumps to the initial state.
[0115] In addition, if the communication device is a first service board of a frame-type network device, the network processor can also send a broadcast message indicating the fault of the first port to the second service board through the network board in the fault confirmation state of the state machine.
[0116] Next, the rapid fault detection and path switching process of box-type network devices is introduced.
[0117] See also Figure 5 , which is a schematic diagram of a process of rapid fault detection and path switching provided in an embodiment of the present application. Figure 5 The process shown may include the following S201-S205.
[0118] S201: Perform rapid fault detection.
[0119] In an embodiment of the present application, the network processor may perform rapid fault detection for one or more ports included in the network device. Taking the first port as an example, the network processor may count the traffic received on the physical bus before parsing the traffic to obtain a statistical result, and determine whether the first port is faulty based on the statistical result. Regarding how to determine whether the first port is faulty based on the statistical result, reference may be made to the relevant description in the previous text, and no repeated description is made here.
[0120] S202: Determine whether a fault is detected.
[0121] If a fault is detected, execute S203, otherwise, continue to execute S201.
[0122] In an example, the network processor may determine that the first port is faulty when the traffic counted in each of a plurality of consecutive statistical periods is less than a first traffic threshold.
[0123] In one example, after determining that the first port is faulty, fault indication information of the first port may be saved.
[0124] S203: Perform path switching.
[0125] In one example, if the network processor receives a first service message and determines that the egress port for forwarding the first service message includes a first port and a second port, the network processor further quickly switches the first service message to the second port for forwarding according to the fault indication information of the first port.
[0126] S204: Check whether there is a false detection.
[0127] If it is determined that the first port failure is a false detection, then S205 is executed: switching the original path. After S205 is executed, S201 is continued to be executed.
[0128] If it is determined that the first port failure is not a false detection, then continue to execute S201.
[0129] Among them, the specific implementation method of checking whether it is a false detection is: when the storage time of the fault indication information of the first port reaches a preset time, the fault indication information is cleared. In one example, if the central processor determines that the state of the first port is up, it means that the fault of the first port is determined to be a false detection. In this case, since the fault indication information is cleared, the network processor can continue to use the first port to forward service messages, thereby realizing rapid path switching. If the central processor determines that the state of the first port is down, it means that the fault of the first port is determined not to be a false detection. In this case, the central processor will delete the forwarding table entry corresponding to the first port in the first forwarding table entry to obtain the second forwarding table entry. Therefore, the fault indication information is cleared, and the service message will not be forwarded through the faulty first port.
[0130] Next, the rapid fault detection and path switching process performed by the first service board of the frame-type network device is introduced.
[0131] See also Figure 6 , this figure is a schematic diagram of another process of rapid fault detection and path switching provided in an embodiment of the present application. Figure 6 The process shown may include the following S301-S307.
[0132] S301: Execute rapid fault detection.
[0133] S302: Determine whether a fault is detected.
[0134] If a fault is detected, execute S303 , otherwise, continue to execute S301 .
[0135] Regarding S301-S302, its specific implementation is the same as S201-S202. For related content, please refer to the description of S201-S202 in the previous text, and no repeated description will be made here.
[0136] S303: Announcement of whole equipment failure.
[0137] In one example, if the network processor of the first service board determines that the first port is faulty, a broadcast message indicating the fault of the first port may be sent to at least one second service board through the network board.
[0138] S304: Determine whether a fault notification is received.
[0139] In one example, the second business board may also execute steps S301-S303. In other words, if the network processor of the second business board determines that the third port is faulty, a broadcast message (i.e., a fault notification) indicating the third port fault may be sent to the first business board through the network board.
[0140] If a fault notification is received, S305 is executed. In addition, after S303 is executed, S305 is also executed.
[0141] S305: Perform path switching.
[0142] In one example, if the network processor of the first service board determines that the first port is faulty, then if the network processor receives the first service message and determines that the egress port for forwarding the first service message includes the first port and the second port, the network processor of the first service board can quickly switch the first service message to the second port for forwarding.
[0143] In yet another example, if the first service board receives a broadcast message indicating a failure of the third port, the network processor of the first service board may send the service message originally sent to the third port to other ports, thereby achieving fast path switching.
[0144] S306: Check whether there is a false detection.
[0145] The checking whether there is a false detection mentioned here refers to checking S302 to determine whether a fault is detected or not.
[0146] If it is determined that the first port failure is a false detection, then S307 is executed: the original path is switched. After S307 is executed, S301 and S304 are continued to be executed.
[0147] If it is determined that the first port failure is not a false detection, then continue to execute S301 and S304.
[0148] Based on the fault determination method provided in the above embodiments, the embodiments of the present application also provide a corresponding network processor. Next, the network processor is introduced in conjunction with the accompanying drawings.
[0149] See also Figure 7 , which is a schematic diagram of the structure of a network processor provided in an embodiment of the present application. Figure 7The network processor 700 shown may belong to a communication device, and the network processor 700 includes: a processing unit 701 .
[0150] The processing unit 701 is used to count the traffic received on the physical bus to obtain statistical results before parsing the traffic, and the communication device includes the physical bus and a first port, and the physical bus is used to connect the first port and the network processor; according to the statistical results, it is determined that the first port is faulty.
[0151] In a possible implementation, the statistical result includes statistical results corresponding to each statistical period in multiple consecutive statistical periods, and the processing unit 701 is used to: determine that the first port is faulty when the traffic counted in each statistical period is less than a first traffic threshold.
[0152] In a possible implementation, the processing unit 701 is used to: run a state machine when a rapid fault detection function is enabled; count the traffic in a fault detection state of the state machine; and determine the first port fault based on the statistical result in the fault detection state.
[0153] In a possible implementation manner, after determining that the first port is faulty, the state of the state machine is switched from the fault detection state to the fault confirmation state.
[0154] In a possible implementation, the processing unit 701 is further configured to: before counting the traffic, determine whether the traffic received by the physical bus within a preset time period is greater than a second traffic threshold.
[0155] In one possible implementation, the determining that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold includes: in the initial state of the state machine, determining that the traffic received by the physical bus within a preset time period is greater than the second traffic threshold; wherein: after determining in the initial state that the traffic received by the physical bus within the preset time period is greater than the second traffic threshold, the state of the state machine is switched from the initial state to the fault detection state.
[0156] In a possible implementation manner, the processing unit 701 is further configured to: save fault indication information indicating a fault of the first port.
[0157] In a possible implementation, the network processor further includes: a receiving unit 702 and a sending unit 703;
[0158] The receiving unit 702 is used to receive a first business message; the processing unit 701 is also used to determine, based on a first forwarding table, that the egress port for forwarding the first business message includes the first port and the second port, and the communication device includes the second port; the sending unit 703 is used to select, based on the fault indication information, to forward the first business message through the second port.
[0159] In a possible implementation, the processing unit 701 is further used to: clear the fault indication information when the storage time of the fault indication information reaches a preset time, and the preset time is greater than the time for the central processor of the communication device to determine the port status of the first port.
[0160] In one possible implementation, after determining that the first port fault occurs, when the state of the state machine is switched from the fault detection state of the state machine to the fault confirmation state, the processing unit 701 saves the fault indication information in the fault confirmation state, and when the storage time of the fault indication information reaches the preset time, the state of the state machine is switched from the fault confirmation state to the initial state of the state machine.
[0161] In one possible implementation, after clearing the fault indication information, if the central processor determines that the state of the first port is closed, the network processor includes a receiving unit 702, which is also used to: receive a second service message; the processing unit 701 is also used to determine, based on a second forwarding table, that the output port for forwarding the second service message includes the second port, wherein the second forwarding table is a forwarding table obtained by the central processor by updating the first forwarding table when determining that the state of the first port is closed; the network processor includes a sending unit 703, which is also used to forward the second service message through the second port.
[0162] In one possible implementation, after clearing the fault indication information, if the central processor determines that the status of the first port is on, the network processor includes a receiving unit 702, which is also used to: receive a third business message; the processing unit 701 is also used to determine, based on the first forwarding table, that the egress port for forwarding the third business message includes the first port and the second port; the network processor includes a sending unit 703, which is also used to select to forward the second business message through the first port.
[0163] In a possible implementation manner, the communication device is a box-type network device.
[0164] In a possible implementation manner, the communication device is a first service board in a frame-type network device.
[0165] In one possible implementation, the frame network device also includes a second business board and a network board, the network board is connected to each business board in the frame network device, and the each business board includes the first business board and the second business board. The network processor includes a sending unit 703, which is also used to: send a broadcast message to the second business board through the network board, and the broadcast message indicates a failure of the first port so that the second business board can perform path switching based on the broadcast message.
[0166] The present application also provides a device, see Figure 8 As shown, the device 800 includes: a processor 810, a communication interface 820 and a memory 830. The number of the processor 810 in the device 800 can be one or more. Figure 8 In the embodiment of the present application, the processor 810, the communication interface 820 and the memory 830 may be connected via a bus system or other means, wherein: Figure 8 The connection via bus system 840 is taken as an example.
[0167] The processor 810 may be a central processing unit (CPU), a network processor (NP) or a combination of a CPU and a NP. The processor 810 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD) or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL) or any combination thereof.
[0168] The memory 830 may include a volatile memory (English: volatile memory), such as a random-access memory (RAM); the memory 830 may also include a non-volatile memory (English: non-volatile memory), such as a flash memory (English: flash memory), a hard disk drive (HDD) or a solid-state drive (SSD); the memory 830 may also include a combination of the above-mentioned types of memory. The memory 830 may, for example, store statistical results of traffic statistics in at least one statistical period.
[0169] Optionally, the memory 830 stores an operating system and a program, an executable module or a data structure, or a subset thereof, or an extended set thereof, wherein the program may include various operating instructions for implementing various operations. The operating system may include various system programs for implementing various basic services and processing hardware-based tasks. The processor 810 may read the program in the memory 830 to implement the fault determination method provided in the embodiment of the present application.
[0170] The bus system 840 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus system 840 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 8 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0171] An embodiment of the present application also provides a computer-readable storage medium, including instructions or computer programs, which, when executed on a computer, enable the computer to execute the fault determination method provided in the above embodiment.
[0172] The embodiments of the present application also provide a computer program product including instructions or a computer program, which, when executed on a computer, enables the computer to execute the fault determination method provided in the above embodiments.
[0173] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0174] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0175] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units is only a logical business division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.
[0176] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0177] In addition, each business unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software business units.
[0178] If the integrated unit is implemented in the form of a software business unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0179] Those skilled in the art will appreciate that in one or more of the above examples, the services described in the present invention may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, the services may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any media that facilitates the transmission of computer programs from one place to another. Storage media may be any available media that can be accessed by a general-purpose or special-purpose computer.
[0180] The above specific implementation modes further describe the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above are only specific implementation modes of the present invention.
[0181] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person of ordinary skill in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features thereof may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. A fault determination method, characterized in that: Applied to a communication device, the communication device includes a network processor, and the method includes: The network processor counts the traffic received on the physical bus to obtain a statistical result before parsing the traffic received on the physical bus, wherein the communication device includes the physical bus and a first port, and the physical bus is used to connect the first port and the network processor; The network processor determines that the first port is faulty according to the statistical result.
2. The method according to claim 1, characterized in that The statistical result includes a statistical result corresponding to each statistical period in a plurality of consecutive statistical periods, and the network processor determines the first port fault according to the statistical result, including: The network processor determines that the first port is faulty when the traffic counted in each statistical period is less than a first traffic threshold.
3. The method according to claim 1 or 2, characterized in that: The counting of the traffic includes: The network processor runs a state machine when the rapid fault detection function is turned on; The network processor counts the traffic in a fault detection state of the state machine; The network processor determines, according to the statistical result, that the first port is faulty, including: The network processor determines, in the fault detection state, that the first port is faulty according to the statistical result.
4. The method according to claim 3, characterized in that: After the network processor determines that the first port is faulty, the state of the state machine is switched from the fault detection state to the fault confirmation state.
5. The method according to claim 3 or 4, characterized in that: Before counting the traffic, the method further includes: The network processor determines that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold.
6. The method according to claim 5, characterized in that The network processor determines that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold, including: The network processor determines, in the initial state of the state machine, that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold; wherein: After the network processor determines in the initial state that the traffic received by the physical bus within a preset time period is greater than a second traffic threshold, the state of the state machine switches from the initial state to the fault detection state.
7. The method according to any one of claims 1 to 6, characterized in that: The method further comprises: The network processor stores fault indication information indicating a fault of the first port.
8. The method according to claim 7, characterized in that The method further comprises: The network processor receives a first service message, and determines, according to a first forwarding table, that an egress port for forwarding the first service message includes the first port and a second port, and the communication device includes the second port; The network processor selects, according to the fault indication information, to forward the first service message through the second port.
9. The method according to claim 7 or 8, characterized in that: The method further comprises: The network processor clears the fault indication information when the storage time of the fault indication information reaches a preset time, and the preset time is greater than the time for the central processor of the communication device to determine the port status of the first port.
10. The method according to claim 9, characterized in that After the network processor determines that the first port is faulty, when the state of the state machine is switched from the fault detection state of the state machine to the fault confirmation state, the network processor saves the fault indication information in the fault confirmation state, and when the storage time of the fault indication information in the network processor reaches the preset time, the state of the state machine is switched from the fault confirmation state to the initial state of the state machine.
11. The method according to claim 9 or 10, characterized in that: After clearing the fault indication information, if the central processor determines that the state of the first port is closed, the method further includes: The network processor receives a second service message, and determines, according to a second forwarding table, that an egress port for forwarding the second service message includes the second port, wherein the second forwarding table is a forwarding table obtained by the central processor by updating the first forwarding table when determining that the state of the first port is closed; The network processor forwards the second service message through the second port.
12. The method according to claim 9 or 10, characterized in that: After clearing the fault indication information, if the central processor determines that the state of the first port is open, the method further includes: The network processor receives a third service message, and determines, according to the first forwarding table, that an egress port for forwarding the third service message includes the first port and the second port; The network processor selects to forward the second service message through the first port.
13. The method according to any one of claims 1 to 12, characterized in that: The communication device is a box-type network device.
14. The method according to any one of claims 1 to 12, characterized in that: The communication device is a first service board in a frame-type network device.
15. The method according to claim 14, characterized in that The frame-type network device further includes a second service board and a network board, the network board is connected to each service board in the frame-type network device, and each service board includes the first service board and the second service board. The method further includes: The network processor sends a broadcast message to the second service board through the network board, where the broadcast message indicates a failure of the first port, so that the second service board performs path switching based on the broadcast message.
16. A network processor, characterized in that: The network processor belongs to a communication device, and the network processor comprises: a processing unit; The processing unit is used to count the traffic received on the physical bus to obtain statistical results before parsing the traffic, and the communication device includes the physical bus and a first port, and the physical bus is used to connect the first port and the network processor; according to the statistical results, it is determined that the first port is faulty.
17. A device, characterized in that include: Processor and memory; The memory is used to store instructions or computer programs; The processor is used to execute the instructions or computer programs to perform the method according to any one of claims 1 to 15.
18. A computer-readable storage medium, characterized in that: The method comprises instructions or computer programs, which, when executed on a computer, enable the computer to execute the method according to any one of claims 1 to 15.