Device failure handling method and device, storage medium, and electronic device

By transmitting fault information to other BMC storages in the wireless LAN through the BMC's coprocessor, the problem of difficulty in repairing faults after a BMC abnormal hang is solved, and rapid fault information acquisition is achieved, reducing the difficulty of equipment maintenance.

CN118869438BActive Publication Date: 2025-09-09INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411170272.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-23
Publication Date
2025-09-09
Estimated Expiration
2044-08-23

AI Technical Summary

Technical Problem

In the prior art, when a BMC crashes abnormally, it is difficult to repair the fault and quickly obtain fault information, resulting in high difficulty in equipment maintenance.

Method used

By adding the current baseboard management controller to a designated wireless LAN through the coprocessor of the BMC, fault information is obtained and transmitted to other BMCs for storage, thereby enabling rapid sharing of fault information.

Benefits of technology

In the event of a serious fault such as an unrecoverable BMC hang, fault information can be quickly obtained without complex acquisition methods, reducing the difficulty of fault repair, improving the speed of fault repair, and reducing the difficulty of equipment maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118869438B_ABST
    Figure CN118869438B_ABST
Patent Text Reader

Abstract

An embodiment of the present application provides a device fault processing method and apparatus, a storage medium, and an electronic device, wherein the method includes: when the current baseboard management controller of the current physical device has not joined the wireless local area network, adding the current baseboard management controller to a designated wireless local area network, wherein the designated wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; when a fault occurs in the current baseboard management controller, obtaining target fault information corresponding to the fault occurring in the current baseboard management controller through a coprocessor, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; and sending the target fault information to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computers, and more specifically, to a method and apparatus for handling device failures, a storage medium, and an electronic device. Background Art

[0002] A Baseboard Management Controller (BMC) is a specialized controller used to monitor and manage physical devices (such as servers and network switches). It provides a system-level way to manage and maintain device hardware. If a BMC crashes, it can cause physical device monitoring to become unavailable, potentially impacting services.

[0003] In related technologies, a remote access and control function can be configured on a physical device. Through this function, the console of the physical device can be accessed to observe the startup process or system logs, and abnormal information before the BMC hangs can be found, so that the cause of the abnormality can be found in time after the BMC abnormality occurs.

[0004] However, the above-mentioned device fault handling method is difficult to repair when a serious fault such as an unrecoverable BMC hang occurs, or when there is insufficient means to obtain logs on site, which increases the difficulty of equipment maintenance. Therefore, it can be seen that the server fault handling method in the related art has the problem of difficulty in fault repair. Summary of the Invention

[0005] The present application provides a device fault handling method and apparatus, a storage medium, and an electronic device to at least solve the problem that the server fault handling method in the related art is difficult to repair the fault.

[0006] According to one aspect of an embodiment of the present application, a device fault handling method is provided, comprising: in a case where a current baseboard management controller of a current physical device has not joined a wireless local area network, adding the current baseboard management controller to a designated wireless local area network through a coprocessor of the current baseboard management controller, wherein the designated wireless local area network is a wireless local area network obtained by networking a group of baseboard management controllers of physical devices as nodes; in a case where a fault occurs in the current baseboard management controller, obtaining target fault information corresponding to the fault occurring in the current baseboard management controller through the coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; and sending the target fault information to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

[0007] According to another aspect of an embodiment of the present application, a device fault handling device is also provided, including: a joining unit, for joining the current baseboard management controller of the current physical device to a designated wireless local area network through a coprocessor of the current baseboard management controller when the current baseboard management controller has not joined the wireless local area network, wherein the designated wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; an acquisition unit, for acquiring target fault information corresponding to the fault occurring in the current baseboard management controller through the coprocessor of the current baseboard management controller when a fault occurs in the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; a first sending unit, for sending the target fault information to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

[0008] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored. When the computer program is executed by a processor, the method steps in any of the above examples are implemented.

[0009] According to another aspect of an embodiment of the present application, an electronic device is also provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor implements the method steps in any of the above examples when executing the computer program.

[0010] In an embodiment of the present application, a method is adopted in which fault information is transmitted to other physical devices for storage through a local area network when a BMC is abnormal. When the current baseboard management controller of the current physical device has not joined the wireless local area network, the current baseboard management controller is added to a specified wireless local area network, wherein the specified wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; when a fault occurs in the current baseboard management controller, the target fault information corresponding to the fault occurring in the current baseboard management controller is obtained through the coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; the target fault information is transferred to the target baseboard management controller. The BMC is then sent to other baseboard management controllers other than the current baseboard management controller in the specified wireless local area network, so that the target fault information is stored in the physical devices to which the other baseboard management controllers belong. When a BMC fails, the fault information is obtained through the coprocessor, and the obtained fault information is transmitted to the physical devices to which other BMCs belong in the same local area network for storage. Therefore, when a serious fault such as an unrecoverable hang occurs in the BMC, the fault information of the BMC can be quickly obtained without complicated acquisition means, which can achieve the purpose of reducing the difficulty of fault repair, achieve the technical effect of increasing the fault repair speed and reducing the difficulty of equipment maintenance, and thus solve the problem of fault repair difficulty in the server fault handling method in the related technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a hardware structure block diagram of an optional server device according to an embodiment of the present application;

[0012] Figure 2 This is a flowchart of an optional device failure handling method according to an embodiment of the present application;

[0013] Figure 3 is a flowchart of another optional device failure handling method according to an embodiment of the present application;

[0014] Figure 4 is a schematic diagram of an optional device failure handling method according to an embodiment of the present application;

[0015] Figure 5 is a flowchart of another optional device failure handling method according to an embodiment of the present application;

[0016] Figure 6 This is a structural block diagram of an optional device fault handling apparatus according to an embodiment of the present application;

[0017] Figure 7 This is a structural block diagram of a computer system of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0018] The embodiments of the present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence.

[0020] The method embodiments provided in the embodiments of the present application can be executed in a server device or a similar computing device. Taking running on a server device as an example, Figure 1 This is a hardware structure diagram of an optional server device according to an embodiment of the present application. Figure 1 As shown, the server device may include one or more ( Figure 1 Only one is shown in the figure) a processor 102 (the processor 102 may include but is not limited to a processing device such as an MCU (Microcontroller Unit) or an FPGA (Field-Programmable Gate Array)) and a memory 104 for storing data. The server device may also include a transmission device 106 and an input / output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above server device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0021] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the device fault handling method in the embodiment of the present application. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implementing the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories may be connected to the server device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0022] The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by a communication provider of the server device. In one embodiment, the transmission device 106 includes a NIC (Network Interface Controller), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be an RF (Radio Frequency) module, which is used to communicate with the Internet wirelessly.

[0023] According to one aspect of an embodiment of the present application, a method for handling a device failure is provided. Figure 2 FIG. 1 is a flow chart of an optional device failure handling method according to an embodiment of the present application, such as Figure 2 As shown, the process includes the following steps:

[0024] Step S202: If the current baseboard management controller of the current physical device has not joined a wireless local area network, join the current baseboard management controller to a designated wireless local area network, wherein the designated wireless local area network is a wireless local area network formed by networking a group of baseboard management controllers of physical devices as nodes;

[0025] Step S204, when a fault occurs in the current baseboard management controller, obtaining target fault information corresponding to the fault of the current baseboard management controller through a coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault of the current baseboard management controller;

[0026] Step S206 : Send the target fault information to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

[0027] The device failure handling method in this embodiment can be applied to the scenario of handling BMC failures of physical devices. BMC is a dedicated controller for monitoring and managing physical devices. It has its own processor and memory. Physical devices can be servers, physical switches, etc. In this embodiment, the server is used as an example for explanation. BMC provides a hardware monitoring and management method independent of the operating system, ensuring the stability and reliability of the server. With the continuous upgrading of server equipment and the continuous expansion of the scale of data centers, the role of BMC has become more and more prominent. Once the BMC has an abnormal hang (system crashes), it will cause the server to be unable to monitor and even affect the business. Therefore, it is an important task of BMC management to promptly find the cause of the abnormality and repair it after the BMC has an abnormality.

[0028] In related technologies, after a BMC crashes, the following methods can be used to troubleshoot and repair the abnormality: configure the remote access and control function on the server, access the server console through this function, observe the startup process or system logs, and look for abnormal information before the BMC crashes, so as to find the cause of the abnormality in time after the BMC crashes.

[0029] For example, if the server is configured with KVM (Keyboard Video Mouse) over IP (Internet Protocol) function, you can access the server console through KVM over IP, observe the startup process or system log, and look for abnormal information before the BMC hangs.

[0030] In addition, troubleshooting and repairing a BMC crash may also include: checking the BMC log. If the BMC can log in normally again and recorded logs before the crash, these logs may contain important clues to the cause of the fault; troubleshooting hardware faults. Based on the server configuration changes before the fault, check the BMC-related hardware connections, such as the SPI (Serial Peripheral Interface) flash memory, I2C (Inter-Integrated Circuit) bus, and power supply, to ensure that there are no loose or damaged components; using external debugging tools. For deeper problems, use the JTAG (Joint Test Action Group) or SWD (Serial Wire Debug) debugging interface and connect to a professional debugger (such as J-Link) for more detailed hardware status inspection and fault location; dump (export) the image in the BMC of the faulty machine and refresh it to a normal machine to check whether the fault phenomenon exists. If so, check the corresponding configuration file for abnormalities and speculate the cause of the fault based on the above information.

[0031] However, the aforementioned BMC health management methods are limited to scenarios where BMC logs or images are available. Analysis is performed only after obtaining the BMC logs or images. If a machine experiences a serious problem or logs cannot be obtained on-site, server maintenance becomes more difficult.

[0032] In order to at least partially solve the above technical problems, in this embodiment, when a BMC fails, fault information is obtained through a coprocessor and transmitted to physical devices to which other BMCs belong within the same local area network for storage. Therefore, when a serious fault such as an unrecoverable hang occurs in the BMC, the BMC fault information can be quickly obtained without complex acquisition methods, thereby reducing the difficulty of fault repair, improving the fault repair speed, and reducing the difficulty of equipment maintenance.

[0033] The device fault handling method in this embodiment can be a solution in which multiple BMCs monitor each other's BMC health status. The multiple BMCs can be located in the same computer room or area. To facilitate mutual health status monitoring, multiple BMCs can form a local area network (LAN) within a certain spatial area using a radio frequency wireless communication module. To ensure the security and reliability of data transmission, the LANs that different BMCs are allowed to join can be pre-specified, that is, the BMCs that can form a LAN can be pre-specified.

[0034] Optionally, the device fault handling method in this embodiment can be executed by a control component on a physical device, such as a BMC, a main processor of the BMC, a coprocessor of the BMC, or a combination thereof. Here, the main processor and coprocessor of the BMC refer to different processor roles in the BMC system. The main processor is the core of the BMC and is responsible for executing the main control logic and task processing. The coprocessor assists the main processor in its work and is mainly used to process specific tasks or functions to reduce the burden on the main processor. The key to defining the main processor and coprocessor lies in their responsibilities and functions. The processor roles of different processors can be pre-set as needed, and can also be adjusted according to actual needs.

[0035] For the current physical device, its baseboard management controller is the current baseboard management controller. If the current baseboard management controller is not yet connected to a wireless local area network, the current baseboard management controller can be connected to a designated wireless local area network. The current baseboard management controller can be connected to one or more wireless local area networks, which can be pre-specified. The designated wireless local area network is one of the wireless local area networks that the current baseboard management controller is allowed to join as a node in the wireless local area network. The designated wireless local area network is a wireless local area network formed by networking the baseboard management controllers of a group of physical devices as nodes. Here, the formed wireless local area network refers to a wireless network for information exchange between wireless communication modules.

[0036] A wireless LAN can be established using a wireless communication module. Specifically, a wireless radio frequency module is used to establish a local area network within a specific spatial area. The wireless communication module can be selected based on specific needs, including frequency and model. Considering that machines are relatively concentrated in the same room and the distances between them are relatively short, a 1.2 GHz wireless communication module can be used. This frequency provides greater bandwidth and faster data transmission speeds. For longer distances, a 900 MHz or 443 MHz wireless communication module can be used.

[0037] Optionally, in this embodiment, in order to handle the impact of the fault on the BMC main processor, the functions related to the wireless communication module can be implemented in the coprocessor. The wireless communication module is connected to the coprocessor, and the wireless communication module and the coprocessor can adopt communication methods such as SPI (Serial Peripheral Interface) and UART (Universal Asynchronous Receiver / Transmitter). Correspondingly, adding the current baseboard management controller to the specified wireless local area network can be executed by the coprocessor, and the data transmission and reception functions between local areas are controlled by the coprocessor, which can include but are not limited to at least one of the following: forming a local area network, adding a new node, node deregistration, etc.

[0038] Since it takes some time for a serious fault to become unrecoverable, and the manifestation of the fault varies in different time periods, the cause of the unrecoverable fault can be deduced based on the evolution of the fault. However, the time of occurrence of a serious fault is random (it is usually impossible to determine at which point in time the BMC will become unrecoverable). In order to reduce the risk of missing abnormal information, in this embodiment, if the current baseboard management controller fails, target fault information corresponding to the fault occurring in the current baseboard management controller can be obtained. The target fault information is used to represent the fault occurring in the current baseboard management controller. It can be a description of the fault or identification information of the fault. The description of the fault can be used to describe the location, event, cause, operating status of related components, etc. of the fault. The identification information of the fault can be an identifier of one or more types of fault description information.

[0039] Optionally, the target fault information may correspond to the log information recorded by the current baseboard management controller when the fault occurs, which may be the above-mentioned log information itself, or may be obtained by processing the above-mentioned log information. Taking into account that only part of the log information recorded by the current baseboard management controller may be related to the fault, the log information may be screened to filter out the available log information. In addition, the fault information may also be determined, for example, by combining the operating status of the relevant components to verify whether a fault has actually occurred (i.e., verifying the authenticity of the fault). Fault information whose authenticity is questionable may be ignored, or the fault information may be processed in combination with relevant identification information to distinguish it from fault information that has passed the authenticity verification.

[0040] For the above-mentioned target fault information, the target fault information can be sent to other baseboard management controllers in the specified wireless local area network except the current baseboard management controller. The other baseboard management controllers can be all baseboard management controllers except the current baseboard management controller, or they can be a specified part of the baseboard management controllers, so as to avoid the same fault information being stored on multiple physical devices and occupying storage resources. Here, if each physical device stores the fault information of all other physical devices in the same local area network, it will occupy too many storage resources.

[0041] For other baseboard management controllers, after receiving the above-mentioned target fault information, the received target fault information can be stored, for example, stored in the storage component of this baseboard management controller, or stored in the storage component of the physical device where it is located (located outside this baseboard management controller). The other baseboard management controllers may store identification information and fault information of the baseboard management controllers with corresponding relationships. Optionally, the baseboard management controller can inform itself that it is in normal operating state through heartbeat information or other means. Through the above-mentioned method, the baseboard management controllers can discover irrecoverable faults such as offline from each other, so that they can quickly respond to such faults. If a specified fault is detected in a baseboard management controller (for example, it changes from an online state to an offline state, the transmitted fault information cannot be accurately parsed, etc.), it can package the fault information corresponding to the baseboard management controller and transmit the packaged fault information to the designated device to prompt that the baseboard management controller has a fault.

[0042] Optionally, the execution order of step S202, step S204 and step S206 can be interchanged. For example, step S206 can be executed first, and then step S204. As mentioned above, obtaining fault information and transmitting fault information can be continuously executed actions.

[0043] Through the above steps, when the current baseboard management controller of the current physical device has not joined the wireless local area network, the current baseboard management controller is added to the designated wireless local area network, wherein the designated wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; when a fault occurs in the current baseboard management controller, the target fault information corresponding to the fault occurring in the current baseboard management controller is obtained through the coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; the target fault information is sent to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information, thereby solving the problem of difficulty in fault repair in the server fault handling method in the related art, improving the fault repair speed, and reducing the difficulty of equipment maintenance.

[0044] In an exemplary embodiment, when a current baseboard management controller fails, obtaining target failure information corresponding to the failure of the current baseboard management controller includes:

[0045] S11, when a fault occurs in the current baseboard management controller, obtaining target fault log information shared by the main processor of the current baseboard management controller through the coprocessor;

[0046] S12, searching a designated database using target fault log information via a coprocessor;

[0047] S13, when a target fault identifier corresponding to the target fault log information is found, determining the target fault identifier as the target fault information;

[0048] S14: If no fault identifier corresponding to the target fault log information is found, the target fault log information is determined as the target fault information.

[0049] To reduce the amount of data transmitted within a wireless local area network, fault identifiers can be used to identify fault log information. The correspondence between fault identifiers and fault log information can be pre-agreed upon. By transmitting fault identifiers, the amount of data required for transmission can be reduced while accurately transmitting information, thereby reducing network resource usage and increasing information transmission speed. Optionally, a designated database can be used to record the correspondence between preset fault log information and fault identifiers. By matching the fault log information with the database, the successfully matched fault log information can be converted into a corresponding fault identifier. The fault identifier here can be a fault number or other identification information, which is not limited in this embodiment.

[0050] For the current baseboard management controller, if the current baseboard management controller fails, the main processor can obtain the fault log information recorded by the current baseboard management controller (log information related to the fault recorded by the BMC) and share the obtained target fault log information with the coprocessor, and the coprocessor can obtain the target fault log information shared by the main processor.

[0051] For the acquired target fault log information, the coprocessor can directly use the target fault log information as the target fault information for synchronization within the designated local area network. To reduce the amount of data required for transmission, the coprocessor can use the target fault log information to search a designated database to determine whether there is a fault identifier that matches the target fault log information. If a target fault identifier corresponding to the target fault log information is found, the target fault identifier can be determined as the target fault information, and correspondingly, the target fault identifier is sent to other baseboard management controllers.

[0052] In this embodiment, if the fault identifier corresponding to the target fault log information is not found, the target fault log information can be directly determined as the target fault information, and correspondingly, the target fault log information is sent to other baseboard management controllers. In order to facilitate database maintenance, if the current baseboard management controller is the first node of the designated wireless local area network, it can assign a corresponding fault identifier to the target fault log information and carry the assigned fault identifier in the sent message to synchronize the correspondence between the target fault log information and the assigned fault identifier between the nodes of the designated wireless local area network. After receiving the message sent by the current baseboard management controller, other baseboard management controllers can use the target fault log information and the assigned fault identifier to update their own databases.

[0053] For example, the BMC may be running a Linux system, and developers can implement various applications based on this. When a fault occurs, the BMC will record certain log information and share this part of the log with the coprocessor. The coprocessor matches the above fault log with the database. The database is matched because the logs directly generated by the BMC are relatively large in number of bytes, and it takes a long time to send the message. After matching the existing fault num (number) in the database, the fault num is directly grouped into the message, which can save the message sending time. If the existing fault information in the database cannot be matched, the node will send the BMC fault log to all local area networks as much as possible to facilitate maintenance in the database.

[0054] Optionally, a BMC fault analysis database can be established and gradually learned and improved, so that when a BMC fault occurs, the cause of the fault can be more accurately and quickly determined. This can avoid transmitting various normal BMC power-on and power-off scenarios to other nodes, which may cause false alarms.

[0055] Through this embodiment, by matching the correspondence between the fault log information and the fault identifier in the database, and when the fault identifier is matched, using the matched fault identifier instead of the fault log information to send to other nodes in the local area network, the amount of data required to be transmitted can be reduced and the message sending time can be shortened.

[0056] In an exemplary embodiment, after obtaining, through the coprocessor, target fault log information shared by the main processor of the current baseboard management controller, the method further includes:

[0057] S21, parsing the target fault log information through the coprocessor to obtain a candidate fault object corresponding to the target fault log information;

[0058] S22, detecting the fault state of the candidate fault object by the coprocessor to obtain a fault detection result of the candidate fault object;

[0059] S23, when the fault detection result indicates that the candidate fault object has not failed, determining the first alarm level in a set of alarm levels as the alarm level corresponding to the target fault log information;

[0060] S24 , when the fault detection result indicates that the candidate fault object has a fault, determining the second alarm level in the set of alarm levels as the alarm level corresponding to the target fault log information.

[0061] For the fault log information shared by the main processor, the coprocessor can directly synchronize it or the fault identifier matched to it within the local area network. In order to avoid false alarms of fault information, the coprocessor can identify the fault log information shared by the main processor and conduct a secondary confirmation of the direction of the fault. There are many types of system crashes, and secondary confirmation means that the coprocessor identifies the type based on the log information and then conducts a second confirmation. The input of the secondary confirmation can be the fault log information, and the output is the fault conclusion of the detection.

[0062] For the target fault log information, the coprocessor can parse the target fault log information to obtain a candidate fault object corresponding to the target fault log information. The candidate fault object refers to the object that triggers the current baseboard management controller to record the target fault log information; then the fault state of the candidate fault object is detected to obtain a fault detection result of the candidate fault object. The fault state of the candidate fault object can be determined based on the object state related to the fault cause of the candidate fault object indicated in the target fault log information, that is, based on the object state related to the fault cause of the candidate fault object indicated in the target fault log information, determine whether the candidate fault object has failed, and the obtained fault detection result is used to indicate whether the candidate fault object has failed.

[0063] If the fault detection result indicates that the candidate fault object has not failed, the first alarm level in a set of alarm levels can be determined as the alarm level corresponding to the target fault log information. If the fault detection result indicates that the candidate fault object has failed, the second alarm level in a set of alarm levels can be determined as the alarm level corresponding to the target fault log information, and the second alarm level is higher than the first alarm level. The baseboard management controller can include an alarm level field (lnfo level) in the message synchronizing fault information within the local area network to indicate the alarm level corresponding to the fault information. The higher the alarm level, the greater the probability that the baseboard management controller has failed. Correspondingly, the indication information of the alarm level corresponding to the target fault log information and the target fault information are sent via the same message.

[0064] For example, after reconfirming the fault log and matching it with existing fault information in the database, if a match is found, the message's alarm level can be set to critical (serious alarm). If the BMC experiences an unrecoverable abnormality (unrecoverable hang), maintenance personnel will repair the BMC based on the information recorded by other nodes on the LAN. If the coprocessor does not obtain a consistent result during the second confirmation of the fault, the fault information will also be sent to the LAN at the info level (low priority alarm).

[0065] For example, the coprocessor identifies the type of error based on the log information and then performs secondary confirmation by searching for keywords to identify the direction it points to. If it is identified as a power supply problem, the current power supply voltage can be checked to determine whether there is a problem with the power supply.

[0066] Through this embodiment, by performing a secondary confirmation on the fault log information and setting the corresponding alarm level in the sent message based on whether the secondary confirmed fault conclusions are consistent, subsequent fault tracing can be facilitated. At the same time, the setting of the alarm level can also prompt the possibility of a fault in the baseboard management controller, so that maintenance personnel can discover the fault in time, thereby improving the timeliness of fault discovery and the convenience of fault tracing.

[0067] In an exemplary embodiment, the same check code is used between nodes in the same local area network. When the coprocessor processes a received message, if it parses the check code and finds that it is different from its own check code, it discards the current message, thereby improving message processing efficiency. For a designated wireless local area network, all messages (e.g., heartbeat messages, logoff messages, etc.) sent by nodes in the designated wireless local area network carry the designated check code. Within the designated local area network, different nodes can be distinguished by node ID (ldentifier). To avoid conflicts when multiple nodes in the local area network send and receive data simultaneously, the nodes in the designated wireless local area network send heartbeat messages in sequence within the designated wireless local area network based on the order of the corresponding node IDs. The heartbeat messages sent by the nodes in the designated wireless local area network all carry the designated check code.

[0068] Optionally, the sending of heartbeat messages by nodes in a specified wireless local area network can be executed according to a specified period. Here, the specified period can be determined based on the timestamp or other time information carried in the heartbeat message received from the previous node, or it can be determined according to the agreed message occurrence time. The time for the node to send the heartbeat message can also be specified by other means.

[0069] Correspondingly, when the current baseboard management controller of the current server has not joined the wireless local area network, joining the current baseboard management controller to the designated wireless local area network includes:

[0070] S31, when the current baseboard management controller has not joined the wireless local area network, receiving a heartbeat message in the current space local area network through the wireless communication module of the current server within a specified time period;

[0071] S32, when no message carrying the specified check code is received within the specified time period, determining the current baseboard management controller as the head node of the specified wireless local area network to form a network, thereby obtaining the specified wireless local area network;

[0072] S33, when a heartbeat message carrying a specified check code is received within a specified time period, after receiving a heartbeat message sent by the last node in the specified wireless local area network, a local area network joining application message is sent in the specified wireless local area network to apply to join the specified wireless local area network; when a topology update message is received from the first node of the specified wireless local area network, it is determined that the current baseboard management controller has joined the specified wireless local area network, wherein the topology update message is used to indicate the local area network topology after the specified wireless local area network joins the current baseboard management controller according to the order of the node IDs of the nodes in the specified wireless local area network.

[0073] In this embodiment, if the current baseboard management controller is not connected to a wireless local area network, a heartbeat message from the current spatial local area network is received via the wireless communication module of the current server within a specified time period. The duration of the specified time period is greater than or equal to the duration of the specified period, so that a heartbeat message sent by a node in the specified local area network can be received at least once within the specified time period.

[0074] For example, when the machine in the computer room is powered on for the first time, the coprocessor can receive data in the current space local area network in real time through the wireless communication module during the waiting time to determine whether a wireless local area network already exists. The waiting time can be: the time of the last byte of the dedicated port MAC (Media Access Control) plus a random number. The above time needs to be greater than the time difference between two nodes in a normal local area network sending data. Here, the data sent can be a heartbeat message or heartbeat information.

[0075] If no message carrying the specified check code is received within the specified time period, the current baseboard management controller is determined to be the first node of the specified wireless local area network for networking, thereby obtaining the specified wireless local area network. That is, if no message with the same check code as its own is received, the node is the first node (first node) corresponding to the check code. If a heartbeat message carrying the specified check code is received within the specified time period, the node can apply to join the existing specified wireless local area network. The method of applying to join can be: sending a local area network joining application message within the specified wireless local area network to apply to join the specified wireless local area network.

[0076] Optionally, the LAN joining request message can be sent after receiving the heartbeat message from the last node in the designated wireless LAN. To facilitate new nodes joining the LAN, the first node can wait for a period of time after the last node sends its heartbeat message before sending the next round of heartbeat messages. Sending the request message at this time can reduce message transmission and reception conflicts.

[0077] In a wireless local area network, the head node can maintain and synchronize the local area network topology (the topology of the local area network, which can represent the order between nodes determined based on the node ID). After sending the local area network joining application message, you can wait to receive the topology structure update message sent by the head node of the specified wireless local area network (the topology structure will change after the new node successfully joins); when receiving the topology structure update message sent by the head node of the specified wireless local area network, it can be determined that the current baseboard management controller has joined the specified wireless local area network. The above-mentioned topology structure update message can be used to indicate the local area network topology of the specified wireless local area network after joining the current baseboard management controller, which is updated according to the order of the node IDs of the nodes in the specified wireless local area network.

[0078] For example, if a message with the same checksum as its own is received while waiting for the last byte of the dedicated port's MAC plus a random number, it indicates that the current LAN already exists and only needs to apply to join it. After determining the last node data in the LAN, the node sends a message requesting to join the LAN. Upon receiving this request, the first node in the LAN updates the LAN topology in order of node ID values ​​and sends a message with the function code 0x02. Upon receiving this message, each node in the LAN updates its stored information, including upstream and downstream node IDs, based on the latest topology.

[0079] Here, a wireless communication protocol message may be a message in a data format sent or received between wireless communication modules. When sending data, content is filled in fixed byte positions according to the message protocol format. The receiving end parses the message according to the protocol format and, after making a judgment, takes the next action based on the received content. The message format of the wireless communication protocol message may include a function code. Different function codes can represent specific functions, including but not limited to the function codes shown in Table 1.

[0080] Table 1

[0081] Function code Functional Explanation Ox01 After a single node is powered on, it applies to join the current wireless LAN for the first time 0x02 The first node broadcasts the latest topology of the current LAN Ox03 During normal operation, the node sends heartbeat information 0x04 The node is logged out due to normal restart, power outage, etc. Ox05 When a node fails, it sends a fault message ...... ......

[0082] As can be seen from Table 1, the message with function code 0x02 is the message used by the first node to broadcast the latest topology of the current local area network. Function codes and corresponding functions can be set and modified according to usage requirements. Table 1 is only an example of a function code.

[0083] Through this embodiment, nodes in the same local area network send heartbeat messages in sequence according to the node ID, which can reduce data transmission and reception conflicts while ensuring the timeliness of node status synchronization; within the set waiting time, it is determined whether the wireless local area network already exists based on the check code carried in the received message, and a matching action is performed, which can ensure the accuracy of local area network joining.

[0084] In an exemplary embodiment, to facilitate determination of the data of the last node in a local area network, a message (e.g., a heartbeat message) sent by a node in a designated wireless local area network may carry its own ID and the next node ID (the first containing the ID of the second...the nth containing the ID of the first). For example, the message format of the wireless communication protocol message used in this embodiment may be as shown in Table 2.

[0085] Table 2

[0086] Contents of the Agreement Byte length Example Header 1 0xa5 Check code 2 Ox1 20x34 Self ID 1 Ox02 Next node ID 1 Ox03 Function code 1 Ox01 Alarm level 1 01~03 Data length 1 Ox01 data N Data check code 1 0×FF End of the report 1 0x5a

[0087] Among them, the header and footer can be set by yourself, and the same server manufacturer can set the same content to improve the compatibility between different models; the next node ID refers to the node that sends data next time after the current node sends data, to avoid conflicts when multiple nodes in the local area network send and receive data at the same time; the check code, function code, and alarm level are similar to those mentioned above; data length, data, and data check code are fields related to the data that needs to be sent in the message.

[0088] Correspondingly, the above method also includes: when a heartbeat message carrying a specified check code is received within a specified time period, the received heartbeat message carrying the specified check code is sequentially used as the current heartbeat message to perform a judgment operation until the heartbeat message sent by the last node in the specified wireless local area network is determined, that is, if it is determined that the heartbeat message sent by the last node in the specified wireless local area network has been received, the judgment operation ends; if it is determined that the heartbeat message sent by the last node in the specified wireless local area network has not been received, continue to wait for the heartbeat message sent by the next node.

[0089] The above-mentioned judgment operation can be: comparing the self-ID carried in the current heartbeat message and the next node ID carried in the current heartbeat message; when the self-ID carried in the current heartbeat message is smaller than the next node ID carried in the current heartbeat message, determining that the current heartbeat message is not the heartbeat message sent by the last node in the specified wireless local area network; when the self-ID carried in the current heartbeat message is larger than the next node ID carried in the current heartbeat message, determining that the current heartbeat message is the heartbeat message sent by the last node in the specified wireless local area network.

[0090] Here, since the heartbeat messages are sent in sequence according to the node ID, the current baseboard management controller has not yet joined the specified wireless LAN, and the node ID of the last node cannot be known. By comparing its own ID and the next node ID carried in the current heartbeat message, if the next node ID is smaller than its own ID, it means that the next heartbeat message is sent by the first node, and the heartbeat message received at this time is not the heartbeat message sent by the last node.

[0091] For example, the last message sent by a node can be determined based on the size of the node IDs in the current LAN. The largest node can be compared and saved. Once the ID of the largest node has been sent, it can apply to join the LAN. After receiving the message from the new node applying to join, the first node sorts the message by ID size across the entire LAN and then sends a message. Each node uses this message to determine its own position and remember the previous and next nodes.

[0092] Through this embodiment, the message sent by the node carries its own ID and the next node ID, and based on the size of its own ID and the next node ID carried in the heartbeat message, it is determined whether the heartbeat message of the last node is received, which can improve the convenience of information determination.

[0093] The device failure handling method in the embodiment of the present application is explained below with reference to an optional example. In this optional example, the physical device is a server, and the functions of the function codes in the message are shown in Table 1.

[0094] Figure 3 FIG. 1 is a flow chart of another optional device failure handling method according to an embodiment of the present application, such as Figure 3 As shown, the process of the method includes the following steps:

[0095] Step S302: Receive the message and wait for the time equal to the last byte of the dedicated port MAC plus a random number.

[0096] Step S304: determine whether the verification code in the received message is consistent with the verification code itself. If not, execute step S306; otherwise, execute step S308.

[0097] Step S306: If no message with the same verification code as its own is received, it is determined that the node is the first node (head node) of the local area network corresponding to the verification code;

[0098] Step S308: If a message with the same verification code as its own is received within the waiting time, a message with the function code 0x02 is sent to apply to join the existing LAN;

[0099] Step S310: The head node of the local area network updates the local area network topology and sends the updated local area network topology in the local area network.

[0100] Step S312: Network establishment is completed.

[0101] Through this optional example, within a certain spatial area, multiple BMCs form a local area network through wireless communication modules (the network can use a handshake protocol) so that when an abnormality occurs in a BMC, the coprocessor will send the collected fault information to other nodes in the local area network through the wireless communication module to improve the fault repair rate.

[0102] In an exemplary embodiment, after adding the current baseboard management controller to the designated wireless local area network, the method further includes:

[0103] S41, collecting node information of each node in the designated wireless local area network through the coprocessor, and saving the node information of each node in the designated wireless local area network to a designated configuration file;

[0104] S43, reading the node information of each node in the specified wireless local area network in the specified configuration file through the main processor, and displaying the topology of the specified wireless local area network on the web page interface of the current baseboard management controller based on the read node information of each node in the specified wireless local area network.

[0105] In this embodiment, to facilitate intuitive viewing of the node structure of a specified wireless local area network by relevant personnel, the topology of the specified wireless local area network can be displayed on the web interface of the current baseboard management controller based on the node information of each node in the specified wireless local area network. The displayed topology can include the relationship between nodes and node information of each node, such as the node name (corresponding to the server name), the node location, the node function (the type of service the node is in), etc., and can also include other content that relevant personnel need to know, which is not limited in this embodiment.

[0106] For example, Figure 4 As shown, the wireless local area network includes 5 nodes, which correspond to the baseboard management controllers on 5 servers respectively. The displayed topology structure can include the order between the nodes and the node information, and the node information includes the relevant information of the corresponding server.

[0107] Displaying the topology of a specified wireless local area network on the web interface of the current baseboard management controller can be performed by the main processor. A number of configuration files are stored in the shared flash memory of the main processor and the coprocessor. The designated configuration file can be used to store node information of each node in the specified wireless local area network. The main processor can read the node information of each node in the specified wireless local area network from the designated configuration file and perform the aforementioned display operation based on the read node information of each node in the specified wireless local area network. Optionally, the designated configuration file can also be stored in a location other than the shared flash memory.

[0108] In the scenario where the designated configuration file is stored in the aforementioned public memory, the coprocessor can collect node information of each node in the designated wireless LAN and save the node information of each node in the designated wireless LAN to the designated configuration file. The aforementioned collection process can be performed periodically according to a set collection cycle, can be performed at a set collection time, or can be executed based on an event trigger, which is not limited in this embodiment.

[0109] For example, a configuration file and a data packet are saved in the shared flash, where the configuration file is some configurations of the wireless communication module made by relevant personnel through the main processor and saved in the configuration file. The coprocessor sets the wireless communication module by reading the configuration file; the data packet is some logs that need to be interacted between the main and coprocessors, as well as BMC identity information corresponding to each node ID and other information used for positioning. The data interaction between the main and coprocessors is not limited to this method.

[0110] After personnel set the verification code for the machine's wireless communication module on the BMC web (World Wide Web), the coprocessor reads the code from the configuration file and populates it in the interaction message. The coprocessor stores the collected information about each node in the local area network in the configuration file. The main processor reads the information and displays it on the BMC web page, allowing personnel to visually view the local area network node structure.

[0111] After the network is established, each node can update the topology according to the topology order sent by the first node. After that, each node will send out the machine's identity information once, which is convenient for display and positioning through the web. After that, each node will send out its own heartbeat information one by one to inform itself that the current BMC is operating normally.

[0112] Through this embodiment, the coprocessor saves the collected information of each node in the local area network in a configuration file. After the main processor reads it, it is displayed on the BMC web page, which can facilitate the intuitive display of the local area network node structure and improve the convenience of information viewing.

[0113] In an exemplary embodiment, in some scenarios, a node may be deregistered, where the node can no longer send or receive messages. The node deregistration may include normal deregistration and abnormal deregistration, where normal deregistration may include normal BMC restart and machine power off.

[0114] As an optional implementation, after the current baseboard management controller is added to the specified wireless local area network, when the current baseboard management controller is restarted, a first node deregistration message is sent in the specified wireless local area network, and the first node deregistration message is used to indicate the deregistration of the current baseboard management controller in the specified wireless local area network.

[0115] During maintenance or other scenarios, the BMC can be restarted normally. The BMC will update the operation to the coprocessor, and the coprocessor will send a first node deregistration message, which is used to instruct the current baseboard management controller to deregister from the specified wireless LAN. Combined with Table 2, the above-mentioned first node deregistration message can be a message with a function code of 0×04. After the first node of the designated wireless LAN receives the above-mentioned first node deregistration message, it will update the topology. The update method can be: the first node sends the node ID information of the entire LAN. Here, within a loop, when the first node does not receive a message from a certain node, it will delete the node and update the LAN ID information.

[0116] As another optional implementation, after the current baseboard management controller is added to the designated wireless local area network, when the current baseboard management controller is powered off, the descriptive information describing the loss of AC power to the power supply unit is transmitted to the coprocessor in the form of an interrupt signal; in response to the received interrupt signal, a second node deregistration message is sent within the designated wireless local area network, wherein the second node deregistration message is used to indicate the deregistration of the current baseboard management controller in the designated wireless local area network.

[0117] For the scenario where the current baseboard management controller is powered off, this information can be transmitted to the coprocessor via an interrupt signal. In response to the received interrupt signal, a second node deregistration message is sent within the specified wireless local area network. The second node deregistration message is used to instruct the current baseboard management controller to deregister from the specified wireless local area network. In conjunction with Table 2, the above-mentioned second node deregistration message can be a message with a function code of 0x04. What is transmitted by the above-mentioned interrupt signal can be descriptive information for describing that the power supply unit has lost alternating current (PSU AC LOST, Power Supply Unit Alternating Current LOST). When the machine is powered off, the PSU AC LOST related information can be connected to the coprocessor in the form of an interrupt. Since the machine has a capacitor energy storage function, it is sufficient for the coprocessor to send the message.

[0118] Alternatively, for abnormal node deregistration, when a BMC failure occurs, the coprocessor will send a message with function code 0x05. If the node subsequently sends another message, the head node will update the topology in a manner similar to the previous embodiment. Node deregistration messages triggered by different reasons can be distinguished by different function codes.

[0119] Through this embodiment, when the BMC is restarted normally or the BMC is powered off, a node deregistration message is sent in the wireless LAN, and the LAN topology of the wireless LAN can be updated in time when the node can no longer send or receive messages, which can improve the accuracy of information in the LAN.

[0120] The following describes the device fault handling method in the embodiment of the present application with reference to an optional example. In this optional example, the physical device is a server. In this optional example, a local area network is established in a certain spatial area through a wireless communication module (e.g., a wireless radio frequency module), including scenarios such as server power-on and BMC restart.

[0121] Under normal circumstances, nodes within a local area network (LAN) are interconnected. If a BMC experiences an anomaly, the coprocessor transmits the collected fault information to other nodes within the LAN via wireless communication. This solution can address the issue of BMC crashes during actual use, where pre-crash anomaly logs cannot be retrieved during repair, leading to high troubleshooting costs.

[0122] Figure 5 FIG. 1 is a flow chart illustrating another optional method for secondary confirmation of device failure according to an embodiment of the present application. Figure 5 As shown, the process of the method includes the following steps:

[0123] Step S502: BMC operates normally;

[0124] Step S504: Determine whether the BMC is faulty. If so, collect the fault log and execute step S506. Otherwise, continue to execute step S502.

[0125] Step S506: Share the collected fault logs with the coprocessor to perform secondary confirmation of the fault information;

[0126] Step S508: Based on whether the secondary determination results are consistent, determine whether there is a fault. If so, execute step S510; otherwise, execute step S512;

[0127] Step S510: If the results are inconsistent, the fault information is sent to the local area network at the info level;

[0128] Step S512: If the results are consistent, the fault information is matched with the database, which may be an established BMC fault analysis database;

[0129] Step S514: If the fault num is matched, a message of the matched fault num is sent;

[0130] Step S516: If no fault information in the database is matched, the BMC fault log is sent to the local area network.

[0131] This optional example allows you to send critical BMC fault information from a server to other servers. By collecting logs that may cause BMC abnormalities and sending them to the BMCs of nearby nodes in real time via wireless communication, you can obtain logs during maintenance, facilitate fault tracing after an unrecoverable BMC failure, and improve fault repair efficiency.

[0132] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0133] According to another aspect of the embodiments of the present application, a device fault handling apparatus is provided. This apparatus is used to implement the device fault handling method provided in the above embodiments. Details already described are omitted for clarity. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0134] Figure 6 This is a structural block diagram of an optional device fault handling device according to an embodiment of the present application. Figure 6 As shown, the device includes:

[0135] A joining unit 602 is configured to, if the current baseboard management controller of the current physical device has not joined the wireless local area network, join the current baseboard management controller to a designated wireless local area network through a coprocessor of the current baseboard management controller, wherein the designated wireless local area network is a wireless local area network formed by networking the baseboard management controllers of a group of physical devices as nodes;

[0136] An acquiring unit 604 is configured to acquire, when a fault occurs in the current baseboard management controller, target fault information corresponding to the fault occurring in the current baseboard management controller through a coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller;

[0137] The first sending unit 606 is configured to send the target fault information to other baseboard management controllers except the current baseboard management controller in the designated wireless local area network, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

[0138] It should be noted that the joining unit 602 in this embodiment can be used to execute the above step S202, the acquiring unit 604 in this embodiment can be used to execute the above step S204, and the first sending unit 606 in this embodiment can be used to execute the above step S206.

[0139] Through the embodiments provided by the present application, when the current baseboard management controller of the current physical device has not joined the wireless local area network, the current baseboard management controller is added to the designated wireless local area network through the coprocessor of the current baseboard management controller, wherein the designated wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; when the current baseboard management controller fails, the target fault information corresponding to the fault occurring in the current baseboard management controller is obtained through the coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; the target fault information is sent to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information, thereby solving the problem of difficulty in fault repair in the server fault handling method in the related art, improving the fault repair speed, and reducing the difficulty of equipment maintenance.

[0140] In an exemplary embodiment, the acquisition unit includes: an acquisition module, which is used to obtain target fault log information shared by the main processor of the current baseboard management controller through a coprocessor when a fault occurs in the current baseboard management controller, wherein the target fault log information is the fault log information recorded by the current baseboard management controller; a search module, which is used to use the target fault log information through the coprocessor to search a designated database, wherein the designated database is used to record the correspondence between preset fault log information and fault identifiers; a first confirmation module, which is used to determine the target fault identifier as the target fault information when a target fault identifier corresponding to the target fault log information is found; and a second confirmation module, which is used to determine the target fault log information as the target fault information when a fault identifier corresponding to the target fault log information is not found.

[0141] In an exemplary embodiment, the above-mentioned device also includes: a parsing unit, which is used to parse the target fault log information shared by the main processor of the current baseboard management controller through the coprocessor to obtain a candidate fault object corresponding to the target fault log information; a detection unit, which is used to detect the fault state of the candidate fault object through the coprocessor to obtain a fault detection result of the candidate fault object; a first determination unit, which is used to determine the first alarm level in a group of alarm levels as the alarm level corresponding to the target fault log information when the fault detection result indicates that the candidate fault object has not failed; a second determination unit, which is used to determine the second alarm level in a group of alarm levels as the alarm level corresponding to the target fault log information when the fault detection result indicates that the candidate fault object has failed; wherein the second alarm level is higher than the first alarm level, and the indication information of the alarm level corresponding to the target fault log information and the target fault information are sent through the same message.

[0142] In an exemplary embodiment, the nodes in the designated wireless local area network send heartbeat messages in sequence within the designated wireless local area network according to a specified period based on the order of the corresponding node IDs. The heartbeat messages sent by the nodes in the designated wireless local area network all carry a specified check code, and the specified check code is used to identify that the corresponding heartbeat message is sent by the node in the designated wireless local area network. The joining unit includes: a receiving module, which is used to receive heartbeat messages in the current spatial local area network through the wireless communication module of the current physical device within a specified time period when the current baseboard management controller has not joined the wireless local area network, wherein the duration of the specified time period is greater than or equal to the period duration of the specified period; a third determination module, which is used to determine the current baseboard management controller as the first node of the specified wireless local area network for networking to obtain the specified wireless local area network when no message carrying the specified check code is received within the specified time period; an execution module, which is used to send a local area network joining application message in the specified wireless local area network after receiving the heartbeat message sent by the last node in the specified wireless local area network to apply for joining the specified wireless local area network; and determine that the current baseboard management controller has joined the specified wireless local area network when a topology update message is received from the first node of the specified wireless local area network, wherein the topology update message is used to indicate the local area network topology updated in the order of the node IDs of the nodes in the specified wireless local area network after joining the current baseboard management controller.

[0143] In an exemplary embodiment, a heartbeat message sent by a node in a designated wireless local area network carries its own ID and a next node ID. Correspondingly, the apparatus further includes: a first execution unit configured to, upon receiving a heartbeat message carrying a designated check code within a designated time period, sequentially use the received heartbeat message carrying the designated check code as the current heartbeat message and perform the following judgment operations until the heartbeat message sent by the last node in the designated wireless local area network is determined: comparing the self-ID carried in the current heartbeat message with the next node ID carried in the current heartbeat message; if the self-ID carried in the current heartbeat message is smaller than the next node ID carried in the current heartbeat message, determining that the current heartbeat message is not the heartbeat message sent by the last node in the designated wireless local area network; if the self-ID carried in the current heartbeat message is larger than the next node ID carried in the current heartbeat message, determining that the current heartbeat message is the heartbeat message sent by the last node in the designated wireless local area network.

[0144] In an exemplary embodiment, the above-mentioned device also includes: a second execution unit, which is used to collect node information of each node in the specified wireless local area network through a coprocessor after adding the current baseboard management controller to the specified wireless local area network, and save the node information of each node in the specified wireless local area network to a specified configuration file, wherein the specified configuration file is a configuration file saved in a shared flash memory of the main processor and coprocessor of the current baseboard management controller; a third execution unit, which is used to read the node information of each node in the specified wireless local area network in the specified configuration file through the main processor, and display the topology structure of the specified wireless local area network on the web page interface of the current baseboard management controller based on the read node information of each node in the specified wireless local area network.

[0145] In an exemplary embodiment, the above-mentioned device also includes: a second sending unit, used to send a first node deregistration message in the designated wireless local area network after the current baseboard management controller is added to the designated wireless local area network and when the current baseboard management controller is restarted, wherein the first node deregistration message is used to indicate the deregistration of the current baseboard management controller in the designated wireless local area network; a third execution unit, used to transmit the description information describing the loss of AC power to the power supply unit in the form of an interrupt signal to the coprocessor when the current baseboard management controller is powered off; and in response to the received interrupt signal, sending a second node deregistration message in the designated wireless local area network, wherein the second node deregistration message is used to indicate the deregistration of the current baseboard management controller in the designated wireless local area network.

[0146] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.

[0147] According to another aspect of the embodiments of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium includes a stored program, wherein the program executes the steps of any of the above method embodiments when running.

[0148] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, at least one of the following: a USB flash drive, a RAM (Random Access Memory), a ROM (Read-Only Memory), a mobile hard disk, a magnetic disk, or an optical disk, and other media that can store computer programs.

[0149] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to execute the steps of any of the above method embodiments through the computer program.

[0150] In an exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0151] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.

[0152] According to another aspect of the embodiment of the present application, a computer program product is provided, which includes a computer program / instruction, and the computer program / instruction contains program code for executing the method shown in the flowchart. In such an embodiment, reference is made to Figure 7 The computer program can be downloaded and installed from the network via the communication section 709 and / or installed from the removable medium 711. When the computer program is executed by the central processing unit 701, the various functions provided by the embodiments of the present application are performed. The serial numbers of the embodiments of the present application are for descriptive purposes only and do not represent the merits of the embodiments.

[0153] refer to Figure 7 , Figure 7 This is a structural block diagram of a computer system of an optional electronic device according to an embodiment of the present application. Figure 7 The following schematically shows a block diagram of a computer system structure of an electronic device for implementing an embodiment of the present application. Figure 7As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 into the random access memory (RAM) 703. Various programs and data required for system operation are also stored in the random access memory 703. The CPU 701, the read-only memory 702, and the random access memory 703 are connected to each other via a bus 704. An input / output interface 705 (i.e., an I / O interface) is also connected to the bus 704.

[0154] The following components are connected to the input / output interface 705: an input section 706 including a keyboard, a mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 708 including a hard disk and the like; and a communication section 709 including a network interface card such as a local area network card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like, is installed in the drive 710 as needed so that a computer program read therefrom can be installed into the storage section 708 as needed.

[0155] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from a removable medium 711. When the computer program is executed by the central processing unit 701, the various functions defined in the system of the present application are performed.

[0156] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0157] Obviously, those skilled in the art should understand that the various modules or steps of the above-mentioned embodiments of the present application can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, they can be implemented using program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0158] The above are only preferred embodiments of the present application and are not intended to limit the embodiments of the present application. For those skilled in the art, the embodiments of the present application may be modified and varied in various ways. Any modifications, equivalent replacements, improvements, etc. made within the principles of the embodiments of the present application shall be included in the scope of protection of the embodiments of the present application.

Claims

1. A method for handling equipment failure, characterized in that: include: If the current baseboard management controller of the current physical device has not joined the wireless local area network, adding the current baseboard management controller to a designated wireless local area network, wherein the designated wireless local area network is a wireless local area network formed by networking the baseboard management controllers of a group of physical devices as nodes; In the case where the current baseboard management controller fails, obtaining target fault information corresponding to the fault of the current baseboard management controller through a coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault of the current baseboard management controller; The target fault information is sent to other baseboard management controllers in the designated wireless local area network except the current baseboard management controller, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

2. The method according to claim 1, characterized in that When the current baseboard management controller fails, obtaining target fault information corresponding to the fault of the current baseboard management controller includes: In the case that the current baseboard management controller fails, obtaining target fault log information shared by the main processor of the current baseboard management controller through the coprocessor, wherein the target fault log information is fault log information recorded by the current baseboard management controller; Searching a designated database using the target fault log information by the coprocessor, wherein the designated database is used to record a correspondence between preset fault log information and fault identifiers; In a case where a target fault identifier corresponding to the target fault log information is found, determining the target fault identifier as the target fault information; If no fault identifier corresponding to the target fault log information is found, the target fault log information is determined as the target fault information.

3. The method according to claim 2, characterized in that After acquiring, by the coprocessor, target fault log information shared by the main processor of the current baseboard management controller, the method further includes: parsing the target fault log information by the coprocessor to obtain a candidate fault object corresponding to the target fault log information; detecting the fault state of the candidate fault object by the coprocessor to obtain a fault detection result of the candidate fault object; If the fault detection result indicates that the candidate fault object has not failed, determining a first alarm level in a set of alarm levels as the alarm level corresponding to the target fault log information; In a case where the fault detection result indicates that the candidate fault object has a fault, determining a second alarm level in a set of alarm levels as the alarm level corresponding to the target fault log information; The second alarm level is higher than the first alarm level, and the indication information of the alarm level corresponding to the target fault log information is sent through the same message as the target fault information.

4. The method according to claim 1, wherein The nodes in the designated wireless local area network sequentially send heartbeat messages within the designated wireless local area network according to a specified period based on an order of corresponding node IDs, wherein the heartbeat messages sent by the nodes in the designated wireless local area network each carry a specified check code, and the specified check code is used to identify that the corresponding heartbeat message is sent by a node in the designated wireless local area network; The step of adding the current baseboard management controller of the current physical device to the designated wireless local area network when the current baseboard management controller has not joined the wireless local area network includes: When the current baseboard management controller is not connected to a wireless local area network, receiving a heartbeat message in the current spatial local area network through the wireless communication module of the current physical device within a specified time period, wherein the duration of the specified time period is greater than or equal to the period duration of the specified period; If no message carrying the specified check code is received within the specified time period, the current baseboard management controller is determined as the head node of the specified wireless local area network to form a network, thereby obtaining the specified wireless local area network; In the case of receiving a heartbeat message carrying the specified check code within the specified time period, after receiving the heartbeat message sent by the last node in the specified wireless local area network, sending a local area network joining application message in the specified wireless local area network to apply to join the specified wireless local area network; in the case of receiving a topology structure update message sent by the first node of the specified wireless local area network, determining that the current baseboard management controller has joined the specified wireless local area network, wherein the topology structure update message is used to indicate the local area network topology of the specified wireless local area network after joining the current baseboard management controller according to the order of the node IDs of the nodes in the specified wireless local area network.

5. The method according to claim 4, characterized in that The heartbeat message sent by the node in the designated wireless local area network carries its own ID and the next node ID; The method further comprises: In the case where a heartbeat message carrying the specified check code is received within the specified time period, the following judgment operations are performed on the received heartbeat messages carrying the specified check code as current heartbeat messages in sequence until the heartbeat message sent by the last node in the specified wireless local area network is determined: Compare the own ID carried in the current heartbeat message with the next node ID carried in the current heartbeat message; When the self-ID carried in the current heartbeat message is less than the next node ID carried in the current heartbeat message, determining that the current heartbeat message is not a heartbeat message sent by the last node in the designated wireless local area network; When the self-ID carried in the current heartbeat message is greater than the next node ID carried in the current heartbeat message, it is determined that the current heartbeat message is the heartbeat message sent by the last node in the designated wireless local area network.

6. The method according to claim 1, characterized in that After adding the current baseboard management controller to the designated wireless local area network, the method further includes: collecting node information of each node in the designated wireless local area network through the coprocessor, and saving the node information of each node in the designated wireless local area network to a designated configuration file, wherein the designated configuration file is a configuration file saved in a shared flash memory of a main processor of the current baseboard management controller and the coprocessor; The main processor reads the node information of each node in the specified wireless local area network in the specified configuration file, and based on the read node information of each node in the specified wireless local area network, displays the topology of the specified wireless local area network on the web page interface of the current baseboard management controller.

7. The method according to any one of claims 1 to 6, characterized in that After adding the current baseboard management controller to the designated wireless local area network, the method further includes: When the current baseboard management controller is restarted, sending a first node deregistration message in the designated wireless local area network, wherein the first node deregistration message is used to instruct to deregister the current baseboard management controller in the designated wireless local area network; When the current baseboard management controller is powered off, descriptive information describing the loss of AC power to the power supply unit is transmitted to the coprocessor in the form of an interrupt signal; in response to the received interrupt signal, a second node deregistration message is sent within the designated wireless local area network, wherein the second node deregistration message is used to indicate the deregistration of the current baseboard management controller in the designated wireless local area network.

8. A device for handling equipment failure, characterized in that: include: a joining unit, configured to join the current baseboard management controller of the current physical device to a designated wireless local area network if the current baseboard management controller of the current physical device has not joined the wireless local area network, wherein the designated wireless local area network is a wireless local area network obtained by networking the baseboard management controllers of a group of physical devices as nodes; an acquiring unit, configured to acquire, when a fault occurs in the current baseboard management controller, target fault information corresponding to the fault occurring in the current baseboard management controller through a coprocessor of the current baseboard management controller, wherein the target fault information is used to indicate the fault occurring in the current baseboard management controller; The first sending unit is configured to send the target fault information to other baseboard management controllers except the current baseboard management controller in the designated wireless local area network, so that the physical devices to which the other baseboard management controllers belong store the target fault information.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Control method, substrate management controller and control system

    CN109766110A

  • Equipment management method and device, equipment and storage medium

    CN118075126A