A network device fault repair method, system, electronic device and medium

By acquiring and controlling the status information of network devices, the repair of faulty devices in a multi-controller network can be achieved without affecting the status of other devices, thus solving the problem of network device failures affecting communication capabilities and improving user experience.

CN116155703BActive Publication Date: 2025-11-21ZHENGZHOU YUNHAI INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202310172291.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-27
Publication Date
2025-11-21
Estimated Expiration
2043-02-27

AI Technical Summary

Technical Problem

In a multi-controller network, troubleshooting a network device can affect the network status of other network devices connected to that device, leading to decreased network communication capabilities and a poor user experience.

Method used

By acquiring the status information of network devices, faults can be identified, and other connected network devices can be identified, repaired, and their status controlled to achieve alarm shielding, including methods such as pin reset and power-on reset, to prevent the spread of faults.

Benefits of technology

In a multi-controller network, the repair of any network device failure does not affect the status of other network devices, thus ensuring network communication capabilities and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116155703B_ABST
    Figure CN116155703B_ABST
Patent Text Reader

Abstract

The application discloses a network device fault repairing method and system, electronic equipment and medium, and relates to the technical field of network devices. The method comprises the following steps: acquiring state information of a network device, and judging whether the network device has a fault according to the state information; if the network device has a fault, acquiring other network devices connected with the network device, repairing the network device, and controlling the state of the other network devices to realize alarm shielding. Through the method, when repairing the fault of any network device in a network with multiple controllers, the other network devices connected with the network device can be shielded from alarms, so that the network communication capability of the other network devices is ensured when the network device has a fault. The network device fault repairing system, electronic equipment and medium provided by the application correspond to the method and have the same effect.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of network devices, and in particular to a network device fault repair method and system, an electronic device, and a medium. BACKGROUND

[0002] In the field of server technology, a full network often adopts a multi-controller redundancy mode in hardware design to ensure data reliability and security; in a multi-controller network, controllers and other network devices exchange data through a data exchange chip (GE) for data interaction.

[0003] However, due to the reliability design requirements of storage servers, there is a risk of failure of the full network Internet hardware device; that is, the GE of the controller and other network devices will fail, and when a failure occurs, it needs to be repaired, but in a multi-controller network, the interconnection between network devices will cause the network state of the network device connected to the chip to change, thereby affecting the network communication capability of the related network devices and the user experience.

[0004] Therefore, how to avoid affecting other network devices when repairing a network device is a problem that needs to be solved by those skilled in the art. SUMMARY

[0005] The purpose of the present application is to provide a network device fault repair method and system, an electronic device, and a medium; to solve the problem that when repairing a faulty network device in a multi-controller network, the network state of the network device connected to the network device will be affected, thereby affecting the network communication capability of the related network devices and the user experience.

[0006] To solve the above technical problems, the present application provides a network device fault repair method, comprising:

[0007] Obtaining state information of a network device;

[0008] Determining whether each corresponding network device has failed according to the state information;

[0009] If the network device fails, obtaining other network devices connected to the network device;

[0010] Repairing the network device and controlling the state of the other network devices to achieve alarm shielding.

[0011] Preferably, controlling the state of the other network devices to achieve alarm includes:

[0012] Changing the state of the register in the other network devices from an alarm state to a repair state to achieve alarm shielding.

[0013] Preferably, the network device comprises a controller and a management board.

[0014] If the network device is a controller, the obtaining of other network devices connected with the network device comprises obtaining other controllers connected with the controller.

[0015] If the network device is a management board, the obtaining of other network devices connected with the network device comprises obtaining all controllers connected with the management board.

[0016] Preferably, if the controller fails, the repairing of the network device and the controlling of the state of other network devices to realize alarm shielding comprises:

[0017] repairing the controller and controlling the state of other controllers to realize alarm shielding;

[0018] If the management board fails, the repairing of the network device and the controlling of the state of other network devices to realize alarm shielding comprises:

[0019] obtaining a master controller corresponding to the management board from all controllers;

[0020] repairing the management board through the master controller and controlling the state of all controllers to realize alarm shielding.

[0021] Preferably, the repairing of the network device comprises:

[0022] repairing the network device through pin reset.

[0023] Preferably, if the network device still fails, after repairing the network device through pin reset, the method further comprises:

[0024] repairing the network device through power-on reset.

[0025] Preferably, the method further comprises:

[0026] counting the number of failures of the network device;

[0027] if the number of failures of the network device reaches a preset number within a preset time, reporting an alarm for the network device.

[0028] To solve the above technical problems, the application further provides a network device failure repairing system, comprising:

[0029] a first obtaining module, configured to obtain state information of network devices;

[0030] a judging module, configured to judge whether each corresponding network device fails according to the state information;

[0031] a second obtaining module, configured to obtain other network devices connected to the network device in the case that the judging module judges that the network device is faulty;

[0032] a repairing module, configured to repair the network device and control the state of the other network devices to realize alarm shielding.

[0033] To solve the above technical problems, the present application further provides an electronic device comprising a memory for storing a computer program;

[0034] a processor, configured to execute the computer program to realize the steps of the network device fault repairing method.

[0035] To solve the above technical problems, the present application further provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to realize the steps of the network device fault repairing method.

[0036] The network device fault repairing method provided by the present application comprises: obtaining state information of a network device and judging whether the network device is faulty according to the state information; if the network device is faulty, obtaining other network devices connected to the network device, repairing the network device, and controlling the state of the other network devices to realize alarm shielding. Through the above method, when repairing the fault of any network device in a network with multiple controllers, the other network devices connected to the network device can be shielded from alarms, so that the network communication capability of the other network devices is ensured and the user experience is improved.

[0037] The present application further provides a network device fault repairing system, an electronic device and a computer readable storage medium, which correspond to the network device fault repairing method and have the same beneficial effects. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0039] Figure 1 a connection diagram of controllers in a multi-control server provided by the present application;

[0040] Figure 2 a network interconnection diagram between network devices in a multi-control server provided by the present application;

[0041] Figure 3A flowchart of a network device fault repair method provided by an embodiment of the present application is provided.

[0042] Figure 4 A flowchart corresponding to a network device fault repair method in a specific application scenario provided by an embodiment of the present application is provided.

[0043] Figure 5 A structural diagram of a network device fault repair system provided by an embodiment of the present application is provided.

[0044] Figure 6 A structural diagram of an electronic device provided by an embodiment of the present application is provided. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, any other embodiments obtained by those skilled in the art without creative work fall within the protection scope of the present application.

[0046] The core of the present application is to provide a network device fault repair method, system, electronic device and medium, mainly related to the network device related technical field, mainly applied in a multi-controller network, for solving the problem that when any network device in the network fails, it will not cause the state change of other network devices connected with the network device, avoiding the alarm and other operations of the other network devices, realizing the alarm shielding of the other network devices, and guaranteeing the network communication ability of the other network devices.

[0047] In order to enable those skilled in the art to better understand the present application scheme, the present application will be further described in detail below with reference to the drawings and specific embodiments.

[0048] In the technical field of servers, the storage server hardware design often adopts a multi-controller redundancy mode to ensure data reliability and security. In the latest multi-controller interconnection technology of storage servers, network interconnection is often the preferred option. Different controllers interact and access management data through network interconnection, but due to the reliability design requirements of storage servers, there is a risk of failure of the entire network interconnection network hardware device. The baseboard management controller is the core of server hardware management, and providing reliable and stable network repair capability is an important reliability means for the entire network architecture of the storage server

[0049] Therefore, based on a storage server full network interconnection architecture, Figure 1 A connection diagram of a controller in a multi-controller server is provided. Figure 1The server frame shown in the figure contains a plurality of controllers 1, each of which is redundant backup for the other controllers 1, wherein the controller 1 includes a central processing unit (CPU), a baseboard management controller (BMC) 2 and a complex programmable logic device (CPLD) 3; wherein the CPLD 3 between different controllers 1 is connected through a backplane 4; wherein the backplane 4 is equivalent to a line in the server. Figure 2 A schematic diagram of network interconnection between network devices in a multi-control server is provided; as Figure 2 shown, wherein each controller 1 further includes a data exchange chip (GE) 5 in addition to the devices mentioned above. Wherein the BMC 2 and the CPU are connected with the GE 5 through different ports, and exchange data within the controller 1 through different ports of the GE 5; wherein two multi-control shared management boards 6 devices are also provided; wherein the management board 6 is also provided with a GE 5, and is also provided with a physical layer chip (PHY) 7 and other network devices; wherein the CPU in the above-mentioned controller 1 exchanges data with the management board 6 through the PHY 7, and exchanges data with the GE 5 through a normal network port; the GE 5 on the management board 6 and the GE 5 in the controller 1 are interconnected to form a hardware structure for network interconnection and intercommunication between different controllers 1 in the storage server.

[0050] Based on the above hardware structure, after the server is powered on, the network communication of the controller 1 and the management board 6 is monitored by the BMC 2, when the GE 5 in the controller 1 and the GE 5 on any management board 6 fail, the repair process of the BMC 2 is triggered, in the original way, repairing any GE 5 in the multi-control network interconnection and intercommunication will cause the network state of other related devices to change; in order to solve the above problem, that is, to ensure that the process of repairing the network device will not affect the network communication capability of other controllers 1, and to reduce the possibility of fault propagation, the present application provides a network device fault repair method, Figure 3 A flowchart of a network device fault repair method is provided, as Figure 3 shown, the method specifically includes:

[0051] S10: Obtain the state information of the network device.

[0052] Wherein the network device can be a controller, or a management board connected to the controller; wherein the state information of the network device includes: the fault state of the network device, that is, whether the network device is in a fault state or a normal operating state; and the connection state of the related service port of the network device, wherein the service port is mainly the related port of the GE in the network device, the GE has ports of different links, and is interconnected with different network devices, and the state (LinkDown and LinkUp) of the port is judged by the BMC to determine whether the network communication state is normal.

[0053] The state information can be acquired in real time, or the acquisition frequency can be set, and the state information of the network device can be acquired at a fixed time, and the acquired state information can be stored in a related memory for subsequent use.

[0054] S11: determining whether the corresponding network device has a fault according to the state information.

[0055] In the above step, the state information of the network device is acquired, and in this step, it is determined whether the corresponding network device has a fault according to the related state information. For example, the state of a controller changes from normal to fault, or the port of a GE in the controller changes from LinkUp to LinkDown. The BMC can monitor the state of the network device by setting a network device monitoring module. When the BMC cannot access the network device or detects that the state of the network device is abnormal, subsequent steps such as repairing the network device are needed.

[0056] S12: If the network device has a fault, acquiring other network devices connected to the network device.

[0057] As shown in the above Figure 1 and Figure 2 , the network device and the other network devices have a connection relationship, that is, the controllers are connected through a backplane, and the controller and the management board are connected through the port of the GE. Therefore, when a network device has a fault, the other network devices connected to the network device are first determined, and the other network devices are controlled in the subsequent steps, so that the influence of the faulty network device on the other network devices can be avoided.

[0058] S13: repairing the network device and controlling the state of the other network devices to realize alarm shielding.

[0059] There are various repair methods, and the commonly used methods include a repair method of pin reset of the network device, which mainly triggers a reset action of the hardware reset pin of the network device, and after the hardware reset is completed, the state information of the network device is read again to determine whether the fault is solved. Another method is a repair method of power-on reset of the network device, which specifically sends a power-on reset instruction to the CPLD in the controller through the BMC of the controller, and performs power-on reset operation on the network device through the CPLD. If the InitDone signal of the network device is set, it proves that the power-on reset is successful. The BMC reads the state information again to determine whether the fault is repaired. The above two repair methods are independent of each other, but only one method can be used to repair the fault at a time. The reset method is relatively fast, about 1s, and the power-on reset method is equivalent to restarting the device, which takes a long time.

[0060] The state of the network device can be divided into an alarm state in normal operation, that is, if a relevant fault is detected when the device is running, the alarm information is reported in time, and the device changes from normal operation to standby operation, and when the fault is handled, the device runs again. The repair state in normal operation is also set, that is, after the device detects a relevant fault, no alarm information is reported, and the device continues to run normally. Therefore, when repairing a network device, to avoid causing other network devices to also report alarm information, other network devices need to change from the alarm state to the repair state, so that no alarm information is reported and the running is not stopped, thereby ensuring the network communication capability of the device.

[0061] The network device fault repair method provided in the embodiment can realize alarm shielding of other network devices connected to the network device when repairing the fault of the network device in a network with multiple controllers, avoid causing the fault of the related network devices when the network device fails, ensure the network communication capability of the other network devices, and improve the user experience.

[0062] On the basis of the above embodiment, the state of the other network device is limited to be controlled to realize the alarm, including changing the state of the register in the other network device from the alarm state to the repair state to realize the alarm shielding.

[0063] The register can store specific state information, and the state information is specifically stored by specific data. When the data changes, the corresponding state information of the register also changes.

[0064] The network device includes a CPLD, and the register is arranged in the CPLD. Figure 1 It can be known that the CPLD in the controller is connected through the backboard, so when a certain controller changes from the alarm state to the repair state, the CPLD between the controllers realizes the state synchronization of the register through the backboard, so that the register of the other controller also changes from the alarm state to the repair state.

[0065] The embodiment provides a method for specifically controlling the state of the other network device, which can accurately and effectively control the other network device, avoid the influence of the faulty device, and maintain normal work.

[0066] On the basis of the above embodiment, the network device includes a controller and a management board. Correspondingly, if the network device is the controller, the other network devices connected to the network device include the other controllers connected to the controller. If the network device is the management board, the other network devices connected to the network device include all the controllers connected to the management board.

[0067] The register can store specific state information, and the state information is specifically stored by specific data. When the data changes, the corresponding state information of the register also changes. Figure 2It can be seen that when the controller fails, the management board connected thereto is not affected, and when the controller fails, other controllers connected thereto through the CPLD are generally not affected, but if the controllers are in information interaction, they may be affected, so as to avoid the above-mentioned situation, when the controller fails, other controllers connected thereto are also acquired. When a management board fails, it can be thought that the management board is reset and restarted for repair, and the GE of the management board causes the ports of all controllers connected to the management board to have a Link Down error, so the management board acquires all controllers connected thereto.

[0068] The embodiment defines which other network devices are connected to a specific network device, and guarantees the accuracy of the alarm shielding range.

[0069] On the basis of the above-mentioned embodiment, if the controller fails, the network device is repaired, and the state of the other network device is controlled to realize alarm shielding, which includes repairing the controller and controlling the state of the other controller to realize alarm shielding; if the management board fails, the network device is repaired, and the state of the other network device is controlled to realize alarm shielding, which includes acquiring the main controller corresponding to the management board from all controllers; repairing the management board through the main controller, and controlling the state of all controllers to realize alarm shielding.

[0070] When a network device fails, it needs to be repaired by the BMC. According to the above-mentioned embodiment, other network devices connected to the failed network device are as follows: when the controller fails, the BMC in the controller is used to repair the controller, and the register state of the CPLD in the controller is changed to synchronize the register state of the other network devices, thereby realizing alarm shielding; when the management board fails, no corresponding BMC is set in the management board, so a controller is needed to repair the management board. In theory, the management board can select any controller connected thereto to repair the management board, but in order to facilitate control, the embodiment limits that the management board has a corresponding main controller. When the management board fails, the BMC of the main controller is used to repair the management board, and the register state of the CPLD of the main controller is changed to synchronize the state of the other controllers, thereby realizing alarm shielding.

[0071] The embodiment limits how to repair the network device when the network device of different types fails, and how to realize alarm shielding for other network devices connected thereto. The state of the other network devices can be quickly and accurately controlled to realize alarm shielding.

[0072] On the basis of the above-mentioned embodiment, repairing the network device includes repairing the network device by means of pin reset.

[0073] Among them, the pin reset mode mentioned in the above embodiment is a way that can quickly repair the network device, and through this way, when repairing the controller, it will not affect the state of other controllers connected with the controller, effectively solving the problem of program running away in the controller, and here the specific way of pin reset is not limited, which can be manually reset by the staff, and other types of reset mode are not limited here.

[0074] The embodiment defines a way to repair the network device, which can quickly repair the fault of the network device.

[0075] As a preferred embodiment, the embodiment defines that if the network device still fails after the above repair method, the network device is repaired by the pin reset mode, and then the network device is repaired by the power-on reset mode.

[0076] Among them, the power-on reset is equivalent to re-equipment, which can effectively solve the problem of program running away in the controller, data error in the register and other problems; but the power-on reset is longer than the above-mentioned pin reset, and the equipment will stop working, affecting the work efficiency, but when the above-mentioned pin reset cannot solve the fault, it is necessary to use the power-on reset. If the above method still cannot repair the faulty network device, other repair methods can be used to continue to repair.

[0077] The embodiment defines that the two reset modes are used to repair the faulty network device, which can effectively repair the fault, and first through the pin reset with short time, then through the power-on reset, which can limit the efficiency of repairing the fault.

[0078] On the basis of the above embodiment, it is defined that in addition to the above steps, it also includes: counting the number of network device failures; if the number of network device failures reaches the preset number within the preset time, the network device is alarmed and reported.

[0079] The embodiment considers that if the network device does not report the alarm, the network device may fail after receiving the repair, but still fails soon. At this time, it needs to be considered whether the hardware of the network device is damaged, and the staff needs to check and repair. Therefore, when the network device fails, the failure of the network device needs to be counted, and whether the failure times of the network device reach the preset times within the preset time is judged; effectively preventing a network device from still being wrong after being repaired, avoiding the damage of the related hardware facilities of the network device. In the embodiment, the preset time and the preset times are not limited. If there is only one pin reset mode for repair, because the repair speed is fast, the preset time may be shorter, the times are more, if it is a power-on reset mode, the preset time may be set longer; here the specific can be judged according to the experience of the staff or the setting of the manufacturer or the historical data of the hardware equipment, etc., which is not specifically described here.

[0080] The embodiment limits the need to report the alarm, avoids repeated errors of the network device, and cannot be solved by repair, which needs the staff to check and avoid causing the system to be paralyzed and other problems.

[0081] The application also provides an embodiment of a network device failure repair method in a specific application scenario, in which the network device failure repair method is specifically as shown in Figure 4 The embodiment provides a network device failure repair method in a specific application scenario, which specifically includes the following steps: S20: monitoring state information of a network device; S21: if the state information shows that the network device fails, acquiring other network devices connected with the network device; S22: controlling states of the other network devices by CPLD to realize alarm shielding; S23: repairing the failed network device by a pin reset mode; S24: judging whether the failed network device is repaired successfully; if not, entering S25: repairing the still failed network device by a power-on reset mode; if yes, entering S26: completing network device failure repair; S27: counting failure times of the network device; and S28: if the failure times of the network device reach preset times within a preset time, reporting an alarm of the network device.

[0082] In the embodiment, the default repair mode of power-on reset can complete the repair of the network device. In the step S22, the state of the network device is controlled by the CPLD, specifically by controlling the state of the CPLD control register to control the state of the network device. The step S28 in the embodiment corresponds to the above embodiment, which is not described here.

[0083] The embodiment provides a network device failure repair method in a specific application scenario, which specifically includes the following steps: S20: monitoring state information of a network device; S21: if the state information shows that the network device fails, acquiring other network devices connected with the network device; S22: controlling states of the other network devices by CPLD to realize alarm shielding; S23: repairing the failed network device by a pin reset mode; S24: judging whether the failed network device is repaired successfully; if not, entering S25: repairing the still failed network device by a power-on reset mode; if yes, entering S26: completing network device failure repair; S27: counting failure times of the network device; and S28: if the failure times of the network device reach preset times within a preset time, reporting an alarm of the network device.

[0084] In the above embodiment, the network device fault repairing method is described in detail, and the application further provides a corresponding embodiment of a network device fault repairing apparatus. It should be noted that the application describes the embodiment of the apparatus from two aspects, one is based on the functional module, and the other is based on the hardware.

[0085] Based on the functional module, the application further provides a corresponding embodiment of a network device fault repairing system, as shown in Figure 5 The system comprises:

[0086] A first obtaining module 10 is configured to obtain the state information of the network device.

[0087] A judging module 11 is configured to determine whether each network device corresponding to the state information has a fault.

[0088] A second obtaining module 12 is configured to obtain other network devices connected to the network device when the judging module determines that the network device has a fault.

[0089] A repairing module 13 is configured to repair the network device and control the state of the other network devices to realize alarm shielding.

[0090] Since the embodiment of the network device fault repairing system corresponds to the embodiment of the method, the network device fault repairing system further comprises:

[0091] A first repairing module is configured to repair the network device by means of power-on reset.

[0092] A statistical module is configured to count the number of faults of the network device.

[0093] An alarm module is configured to alarm and report the network device if the number of faults of the network device reaches a preset number within a preset time.

[0094] The network device fault repairing system provided in the embodiment corresponds to the above method, and has the same beneficial effects as the above method.

[0095] Based on the hardware, the application provides an electronic device. Figure 6 The structural diagram of the electronic device provided in the embodiment is shown in Figure 6 The electronic device comprises a memory 20 configured to store a computer program.

[0096] A processor 21 is configured to execute the computer program to realize the steps of the network device fault repairing method mentioned in the above embodiment.

[0097] The electronic device provided by the embodiment can include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, and the like.

[0098] The processor 21 can include one or more processing cores, such as a 4-core processor, an 8-core processor, and the like. The processor 21 can be implemented in at least one of a hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic array (PLA). The processor 21 can also include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a central processing unit (CPU). The coprocessor is a low-power processor for processing data in a standby state. In some embodiments, the processor 21 can be integrated with a graphics processor (GPU) for rendering and drawing content required to be displayed by the display screen. In some embodiments, the processor 21 can further include an artificial intelligence (AI) processor for processing machine learning-related computing operations.

[0099] The memory 20 can include one or more computer-readable storage media, which can be non-transitory. The memory 20 can also include a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash storage devices. In the embodiment, the memory 20 is at least used to store the following computer program 201, wherein the computer program is loaded and executed by the processor 21, and can implement the related steps of the network device fault repair method disclosed in any of the preceding embodiments. In addition, the resources stored by the memory 20 can further include an operating system 202 and data 203, and the storage mode can be temporary storage or permanent storage. The operating system 202 can include Windows, Unix, Linux, and the like. The data 203 can include, but is not limited to, data included in the network device fault repair method, and the like.

[0100] In some embodiments, the electronic device can further include a display screen 22, an input / output interface 23, a communication interface 24, a power supply 25, and a communication bus 26.

[0101] Those skilled in the art can understand that, Figure 6 The structure shown in the figure does not constitute a limitation on the electronic device, and can include more or fewer components than those shown.

[0102] The electronic device provided by the embodiment of the present application comprises a memory and a processor. When the processor executes a program stored in the memory, the following method can be realized: a network device fault repair method.

[0103] The electronic device provided by the embodiment of the present application corresponds to the above method, and has the same beneficial effects as the above method.

[0104] Finally, the present application also provides an embodiment corresponding to a computer readable storage medium. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps recorded in the above method embodiments are realized.

[0105] It can be understood that if the method in the above embodiment is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and executes all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (Read-Only Memory, ROM), a random access memory (Random Access Memory, RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0106] The computer readable storage medium provided by the embodiment of the present application corresponds to the above network device fault repair method, and has the same beneficial effects as the above method.

[0107] The network device fault repair method, system, electronic device and medium provided by the present application are described in detail above. The embodiments in the specification are described in a progressive manner. Each embodiment mainly describes the differences from other embodiments. The same or similar parts of each embodiment can be referred to. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the related parts can be referred to the method part. It should be pointed out that, for ordinary skilled in the art, without departing from the principles of the present application, some improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

[0108] It also needs to be explained that in the present specification, the relational terms such as first and second and the like are used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

Claims

1. A network device fault recovery method, characterized by, The method comprises the steps of: acquiring state information of network devices; the network device comprises a controller and a management board connected with the controller; the controllers are connected through a backboard, and the controller and the management board are connected through a GE port; the state information of the network device comprises fault information of the network device and connection states of relevant service ports of the network device, wherein the service ports are relevant ports of the GE in the network device, the GE has ports with different links, and the ports are interconnected with different network devices to determine whether the network communication state is normal through the state of the port by the BMC; determining whether each corresponding network device has a fault according to the state information; if the network device has a fault, acquiring other network devices connected with the network device; repairing the network device through a pin reset mode, wherein a hardware reset pin of the network device triggers a reset action, and after the hardware reset is completed, the state information of the network device is read again to determine whether the fault is repaired; if the network device still has a fault, the network device is repaired through a power-on reset mode after the network device is repaired through the pin reset mode, wherein the BMC of the controller sends a power-on and power-off instruction to the CPLD in the controller, the CPLD performs a power-on and power-off operation on the network device, and if the InitDone signal of the network device is set, it proves that the power-on and power-off operation is successful; the BMC reads the state information again to determine whether the fault is repaired; when the fault is repaired, the state of the register in the CPLD of the other network devices is changed from an alarm state to a repair state to realize alarm shielding; if the network device is the controller and the controller has a fault, the controller is repaired through the BMC in the controller, and other controllers connected with the controller are acquired; the state of the register in the other controllers is synchronized from an alarm state to a repair state by changing the register state in the CPLD in the controller to realize alarm shielding, wherein the CPLD between the controllers realizes the state synchronization of the register through the backboard; if the network device is the management board and the management board has a fault, all controllers connected with the management board are acquired, and a main controller corresponding to the management board is acquired from the all controllers; the management board is repaired through the BMC of the main controller, and the state of the other controllers is synchronized by changing the register state of the CPLD on the main controller, so that the state of the register in the all controllers is changed from an alarm state to a repair state to realize alarm shielding.

2. The network device failure remediation method of claim 1, wherein, Further comprising: counting the number of faults of the network device; if the number of faults of the network device reaches a preset number within a preset time, the network device is alarmed and reported.

3. A network device fault remediation system, comprising: The method comprises the steps of: a first acquisition module for acquiring state information of network devices; The network device comprises a controller and a management board connected with the controller; the controllers are connected through a backboard, and the controller and the management board are connected through a GE port; the state information of the network device comprises fault information of the network device and connection states of relevant service ports of the network device, wherein the service ports are relevant ports of the GE in the network device, the GE has ports of different links, and the ports are interconnected with different network devices to judge whether the network communication state is normal through the state of the port by the BMC; a judging module for judging whether each corresponding network device appears a fault according to the state information; a second obtaining module for obtaining other network devices connected with the network device in the case that the judging module judges that the network device appears a fault; a repairing module for repairing the network device through a pin reset mode, wherein a hardware reset pin of the network device triggers a reset action, and after the hardware reset is completed, the state information of the network device is read again to judge whether the fault is repaired, if the network device still appears a fault, the network device is repaired through a power-on reset mode after the network device is repaired through the pin reset mode, wherein the BMC of the controller sends a power-on and power-off instruction to the CPLD in the controller, the network device is operated again through power-on and power-off by the CPLD, if the InitDone signal of the network device is set, it proves that the power-on and power-off operation is successful; the BMC reads the state information again to judge whether the fault is repaired; when the fault is repaired, the state of the register in the other network devices is changed from an alarm state to a repair state to realize alarm shielding; if the network device is the controller, the controller appears a fault, the controller is repaired through the BMC in the controller, and other controllers connected with the controller are obtained, the state of the register in the other controllers is synchronized from the alarm state to the repair state through the register state change of the CPLD in the controller to realize alarm shielding, wherein the CPLD between the controllers realizes the state synchronization of the register through the backboard; if the network device is the management board, the management board appears a fault, all controllers connected with the management board are obtained, and a main controller corresponding to the management board is obtained from the all controllers; the management board is repaired through the BMC of the main controller, and the state of the other controllers is synchronized through the register state change of the CPLD on the main controller, so that the state of the register in the all controllers is changed from the alarm state to the repair state to realize alarm shielding.

4. An electronic device, comprising: The network device comprises a memory for storing a computer program; a processor for executing the computer program to realize the steps of the network device fault repairing method in claim 1 or 2.

5. A computer readable storage medium, characterized in that, The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the network device fault repairing method in claim 1 or 2.

Citation Information

Patent Citations

  • Controller and alarm correlation processing method

    CN105790972A

  • Hard disk failure handling method, apparatus, server, and computer-readable medium

    CN109284207A

  • Method, device and equipment for preventing OTN optical channel protection deadlock and storage medium

    CN113114406A