Network management method of management board and electronic equipment

By monitoring the link status, port status, and signal parameters between the management board and storage nodes in real time, and combining this with historical status records, the limitations of network management on the management board in storage devices are resolved. This improves the accuracy of fault location and operational efficiency, ensuring the high reliability and stability of the storage system.

CN121585531APending Publication Date: 2026-02-27INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511710329.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

The network management methods of existing storage devices' management boards have high limitations, making it impossible to effectively monitor multiple nodes sharing a management board with limited hardware resources. This results in inaccurate fault location and low operation and maintenance efficiency, affecting the reliability and stability of the storage system.

Method used

By monitoring the link connection status, network port status, temperature, and signal parameters between the management board and the respective controllers of the storage nodes in real time, a weighted average method is used to determine whether there is a fault in the network link, and historical status records are combined to locate the fault and provide maintenance support.

Benefits of technology

It enables comprehensive monitoring of the network links of the shared management board, improves the accuracy of fault location and operation and maintenance efficiency, and ensures the high reliability and stability of the storage system in complex network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121585531A_ABST
    Figure CN121585531A_ABST
Patent Text Reader

Abstract

The invention discloses a network management method of a management board and electronic equipment, and relates to the technical field of servers, and the method comprises the steps that the server monitors the link connection state, the network port physical state, the temperature and the signal parameter between the management board and the respective controller of at least one storage node in real time; therefore, fault detection of the network link of the shared management board is realized. According to the network management method of the management board provided by the invention, a one-to-one fixed binding mode between the management board and the storage nodes is broken through, comprehensive monitoring on the network state of the multi-node shared management board under limited hardware resources is supported, the limitation of network management on the management board is reduced, and the network management efficiency is improved. And meanwhile, the fault positioning accuracy and the operation and maintenance efficiency are improved, so that high reliability and stability of the storage system are guaranteed in a complex network environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of server technology, and in particular to a network management method and electronic device for a management board. Background Technology

[0002] The management board is the core control component of a storage device. It uses network links to schedule network communication, monitor operational status, and provide local maintenance support for each storage node within the device. In some scenarios, the storage device can manage the network links of the management board.

[0003] In related technologies, storage devices deploy at least one storage node and at least one management board, wherein the at least one storage node and at least one management board are connected one-to-one at the hardware level, meaning that one management board is responsible for managing the storage node connected to it. Specifically, for the network link between each management board and the storage node, the storage device detects whether the network link is normal and generates a prompt message in case of an abnormality, which is used to indicate the network link abnormality. However, the above-mentioned network management method for the management board has significant limitations. Summary of the Invention

[0004] This application provides a network management method and electronic device for a management board, thereby reducing the limitations of network management of the management board.

[0005] This application provides a network management method for a management board, including:

[0006] Determine the link status information of the management board; wherein, the management board is connected to the controller of at least one storage node via network links, and the link status information is used to indicate the connection status between the management board and the controller of at least one storage node.

[0007] The system detects the port status, port temperature, and port signal parameters of the network ports on the management board; wherein, the network port is the port on the management board that connects to the network.

[0008] Based on link status information, port status, port temperature, and port signal parameters, the first state of the management board is determined. The first state is used to indicate whether there is a fault in the network link of the management board at the current moment.

[0009] This application also provides a network management device for a management board, comprising:

[0010] The determination module is used to determine the link status information of the management board; wherein the management board is connected to the controller of at least one storage node through network links, and the link status information is used to indicate the connection status between the management board and the controller of at least one storage node.

[0011] The detection module is used to detect the port status, port temperature, and port signal parameters of the network ports on the management board; wherein, the network port is the port on the management board that is connected to the network;

[0012] The processing module is used to determine the first state of the management board based on link status information, port status, port temperature, and port signal parameters. The first state is used to indicate whether there is a fault in the network link of the management board at the current moment.

[0013] This application also provides an electronic device, including: a memory for storing a computer program; and a processor for executing the computer program to implement the network management method of the management board described above.

[0014] This application also provides a computer-readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the steps of the network management method of the management board described above.

[0015] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the network management method of the management board described above.

[0016] This application enables the server to detect network link faults on the shared management board by real-time monitoring of the link connection status, network port physical status, temperature, and signal parameters between the management board and the controllers of at least one storage node. The network management method for the management board provided in this application breaks through the one-to-one fixed binding mode between the management board and storage nodes, supporting comprehensive monitoring of the network status of multiple nodes sharing a management board under limited hardware resources. This reduces the limitations of network management of the management board, while improving the accuracy of fault location and operational efficiency, thereby ensuring the high reliability and stability of the storage system in complex network environments. Attached Figure Description

[0017] To more clearly illustrate the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram illustrating an application scenario provided in the embodiments of this application;

[0019] Figure 2 A flowchart illustrating a network management method for a management board provided in an embodiment of this application;

[0020] Figure 3A flowchart illustrating the process of determining a first state, provided for an embodiment of this application;

[0021] Figure 4 This is a schematic diagram illustrating a process for updating the status record information of a management board, provided as an embodiment of this application.

[0022] Figure 5 A schematic diagram of the structure of a network management device for a management board provided in an embodiment of this application;

[0023] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Detailed Implementation

[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, any other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of this application.

[0025] It should be noted that, in the description of this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. The terms "first," "second," etc., in this application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0026] First, the technical terms used in this application will be explained.

[0027] Management board: This is the core control component of the storage device. It uses network links to schedule network communication, monitor the operational status, and provide local maintenance support for each storage node within the storage device. The management board integrates a network cable interface, a Universal Serial Bus (USB) interface, and a serial port interface.

[0028] The management board handles data transmission for the management network via a network cable interface, providing a network communication channel for the storage devices. Through USB and serial interfaces, the management board enables firmware upgrades, log export, and low-level debugging, supporting localized operation and maintenance of the storage devices. In some scenarios, the serial interface serves as the core out-of-band management channel, allowing direct access to the underlying hardware to initialize storage device configurations and debug the kernel.

[0029] In some embodiments, the storage device can manage the network of the management board through a centralized management platform, monitor the network status of the management board in real time, and simplify the operation and maintenance process.

[0030] In the era of cloud computing and big data, massive data processing places high demands on the reliability and stability of storage systems. As the core control component of storage devices, the network management stability of the management board is a crucial factor in ensuring the stability of the storage device. Therefore, storage devices can manage the network links of the management board.

[0031] In related technologies, storage devices deploy at least one storage node and at least one management board. To ensure the independence and reliability of the management channel, the at least one storage node and at least one management board are connected one-to-one at the hardware level; that is, one management board is responsible for managing the storage node connected to it. Specifically, for the network link between each management board and the storage node, the storage device detects whether the network link is normal and generates a prompt message in case of an abnormality. The prompt message is used to indicate the abnormality of the network link.

[0032] However, when storage devices are resource-constrained—that is, when the number of management boards that can be deployed in a storage device is less than the number of storage nodes—some storage nodes cannot access the network link of the management board. This results in these storage nodes not receiving normal communication scheduling and status monitoring support, thus rendering the aforementioned network management methods for the management board highly limited. Furthermore, the network management of the management board by the storage device can only provide simple indications of network link failures, without pinpointing the cause of the failure, and lacks a network link control mechanism adapted to multi-storage-node shared scenarios, leading to low overall reliability and stability of the storage device.

[0033] Based on this, this application provides a network management method for a management board. The server monitors in real time the link connection status, network port physical status, temperature, and signal parameters between the management board and the controllers of at least one storage node, thereby enabling fault detection of the shared management board network links. The network management method provided in this application breaks through the one-to-one fixed binding mode between the management board and storage nodes, supporting comprehensive monitoring of the network status of multiple nodes sharing a management board under limited hardware resources. This reduces the limitations of network management of the management board, while improving fault location accuracy and operational efficiency, thus ensuring the high reliability and stability of the storage system in complex network environments.

[0034] To enable those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0035] The specific application environment architecture or hardware architecture upon which the network management method of the management board depends is described here. (References) Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario provided in an embodiment of this application. For example... Figure 1 As shown, it includes a server 11, wherein the server 11 may be, for example, a storage device, and at least one management board and at least one storage node are deployed in the server 11.

[0036] In practical applications, any one of the at least one management boards can be connected to one or more of the at least one storage node via a network link. Server 11 can manage the network of the management board to achieve centralized network communication scheduling, real-time operation status monitoring, and operation and maintenance support for the one or more storage nodes connected to the management board through the network link of the management board.

[0037] like Figure 1 As shown, taking a server with one management board x and two storage nodes (storage node y and storage node z) deployed as an example, the management board x is connected to storage node y and storage node z through network links.

[0038] It should be noted that the execution subject in each embodiment of this application can be a processor, microprocessor, or a device integrating the aforementioned processor or microprocessor, such as a terminal device. The specific execution subject in each embodiment of this application is not limited and can be selected and set according to actual needs. In the following embodiments, a terminal device integrating the aforementioned processor or microprocessor is used as an example for description, which does not constitute a limitation on the actual execution subject.

[0039] It should be noted that, Figure 1 This is merely an example to illustrate one application scenario, and is not intended to limit the application scenario.

[0040] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0041] Figure 2 This is a flowchart illustrating a network management method for a management board provided in an embodiment of this application. Figure 2 As shown, the method may include the following steps:

[0042] S201. Determine the link status information of the management board; wherein, the management board is connected to the controller of at least one storage node through network links, and the link status information is used to indicate the connection status between the management board and the controller of at least one storage node.

[0043] A management board is a component deployed in a server that performs functions such as network communication scheduling, operational status monitoring, and local operation and maintenance support (e.g., firmware upgrades, log export) for storage nodes. In some embodiments, the server may be a storage device, and at least one management board may be deployed in the server. Each of the at least one management board can establish a network connection with the respective controller of at least one storage node. The at least one storage node is a storage node within the server.

[0044] Specifically, the management board includes a switching module, and the controller of the storage node includes a switching module. In some embodiments, for any one of the at least one management board, the switching module of the management board can establish a network connection with the switching module of the respective controller of at least one storage node.

[0045] The link status information of the management board is used to indicate the connection status of the network connection between the switching module of the management board and the switching module of the controller of at least one storage node.

[0046] The server includes a Multiple Controller System (MCS), an Intelligent Platform Management Interface (IPMI), and a Baseboard Management Controller (BMC). In some embodiments, the server's MCS can obtain link status information from the BMC system via the IPMI.

[0047] S202. Detect the port status, port temperature, and port signal parameters of the network port on the management board; where the network port is the port on the management board that connects to the network.

[0048] The management board includes network ports, which are physical interfaces (such as network cable interfaces) specifically designed for accessing and establishing network connections. These network ports are crucial hardware interfaces for data transmission between the management board, the storage node's controller, and external networks, and form the physical foundation for network communication between the management board and the storage node.

[0049] Port status refers to the real-time operating status of a network port. The port status is either normal or abnormal. When the port status is normal, it means that the physical connection of the network port is normal and the data transmission is normal. When the port status is abnormal, it means that the physical connection of the network port is abnormal and / or the data transmission is abnormal.

[0050] In some embodiments, the server's MCS can obtain the port status of the network port from the BMC system via IPMI.

[0051] Port temperature refers to the real-time temperature data of a network port during operation. In some embodiments, the network port of the management board is equipped with a temperature sensor, which the server can use to detect the port temperature of the network port.

[0052] Port signal parameters indicate the signal strength of a network port and are key parameters characterizing the quality of signal transmission from the network port. Port signal parameters may include, for example, signal interface power and signal attenuation. In some embodiments, the port signal parameters of a network port directly affect the stability and effectiveness of data transmission.

[0053] In some embodiments, the server can detect port signal parameters based on a reference signal.

[0054] S203. Based on the link status information, port status, port temperature, and port signal parameters, determine the first state of the management board. The first state is used to indicate whether there is a fault in the network link of the management board at the current moment.

[0055] The first state is the result of the storage device's determination of the operating status of the network link to the management board at the current moment. Its core purpose is to clearly indicate whether there is a fault in the network link between the management board and each storage node controller.

[0056] In some embodiments, the first state can be determined as follows: The server determines link status index values ​​based on link status information, port status index values ​​based on port status, temperature index values ​​based on port temperature and a preset temperature, and signal index values ​​based on port signal parameters and preset parameters. Then, the server performs a weighted average of the link status index values, port status index values, temperature index values, and signal index values ​​according to their respective preset weights to obtain the management board's status index values. The server determines the first state of the management board based on the management board's status index values.

[0057] exist Figure 2In the illustrated embodiment, the server monitors the link connection status, network port physical status, temperature, and signal parameters between the management board and the controllers of at least one storage node in real time, thereby enabling fault detection of the shared management board network link. The network management method for the management board provided in this application breaks through the one-to-one fixed binding mode between the management board and storage nodes, supporting comprehensive monitoring of the network status of multiple nodes sharing a management board under limited hardware resources. This reduces the limitations of network management of the management board, while improving fault location accuracy and operational efficiency, thus ensuring the high reliability and stability of the storage system in complex network environments.

[0058] exist Figure 2 Based on the illustrated embodiment, the following, in conjunction with Figure 3 The method for determining the first state of the management board in the embodiments of this application will be further explained.

[0059] Figure 3 This is a schematic diagram illustrating a process for determining a first state, provided as an embodiment of this application. Figure 3 As shown, the process may include the following steps:

[0060] S301. Determine the connection status of the network link based on the link status information; the connection status is either normal or abnormal.

[0061] The network link connection status is used to indicate the connection status of the network link between the management board and at least one storage node. In some embodiments, if the connection status of the network link between the management board and at least one storage node is normal, the connection status of the network link is determined to be normal; if the connection status of the network link between at least one storage node and the management board is abnormal, the connection status of the network link is determined to be abnormal.

[0062] For example, suppose at least one storage node includes storage node 1, storage node 2 and storage node 3, wherein the network link between the management board and storage node 1 is network link A, the network link between the management board and storage node 2 is network link B, and the network link between the management board and storage node 3 is network link C.

[0063] If the connection status of network link A is normal, the connection status of network link B is normal, and the connection status of network link C is abnormal, then the server determines that the connection status of the network links is abnormal.

[0064] If the connection status of network link A is normal, the connection status of network link B is normal, and the connection status of network link C is normal, then the server determines that the connection status of the network links is normal.

[0065] In some embodiments, the link state information includes a first connection state of the controller of at least one storage node and a second connection state of the management board. The server may determine the connection state of the network link as follows: for each storage node in the at least one storage node, determine whether the first connection state of the storage node's controller is normal; determine whether the second connection state is normal; if the first connection state of the controllers of at least one storage node is normal and the second connection state is normal, determine the connection state as a normal connection state; if the first connection state of the controllers of at least one storage node is abnormal, and / or the second connection state is abnormal, determine the connection state as an abnormal connection state.

[0066] Since the management board is connected to the controller of at least one storage node, for any one of the at least one storage node, the first connection state of the controller of that storage node is used to indicate that the controller of that storage node has detected a connection state between the controller of that storage node and the management board. In some embodiments, if the first connection state of the controller of that storage node is normal, it indicates that the controller of that storage node has detected a normal connection state between the controller of that storage node and the management board; if the first connection state of the controller of that storage node is abnormal, it indicates that the controller of that storage node has detected an abnormal connection state between the controller of that storage node and the management board.

[0067] The second connection status of the management board is used to indicate the connection status detected by the management board with the respective controllers of at least one storage node. If the second connection status is normal, it means that the management board has detected a normal connection status with the respective controllers of at least one storage node; if the second connection status is abnormal, it means that the management board has detected an abnormal connection status with the respective controllers of at least one storage node.

[0068] Therefore, if the first connection status of the controllers of at least one storage node is normal and the second connection status is normal, it indicates that the network link connection status between the management board and at least one storage node is normal. At this time, the server determines that the network link connection status is normal.

[0069] If the first connection state of the controller of at least one storage node is abnormal, and / or the second connection state is abnormal, it indicates that the connection state of the network link between the management board and at least one storage node is abnormal. In this case, the server determines that the connection state of the network link is abnormal.

[0070] S302. Determine the network status of the management board based on the port status, port temperature, and port signal parameters; the network status is either normal or abnormal.

[0071] The network status of the management board is used to indicate the physical connection integrity and data transmission status of the network link corresponding to the network port of the management board at the current moment. In some embodiments, if the network status is normal, it means that the physical connection of the network link corresponding to the network port is complete and the data transmission is normal; if the network status is abnormal, it means that the physical connection of the network link corresponding to the network port is incomplete and / or the data transmission is abnormal.

[0072] In some embodiments, the server may determine the network status as follows: determine whether the temperature of the network port is normal based on the port temperature and a preset temperature range; determine whether the signal of the network port is normal based on the port signal parameters and a preset parameter range; determine the network status as normal if the port status is normal, the network port temperature is normal, and the network port signal is normal; determine the network status as abnormal if at least one of the following exists: the port status is abnormal, the network port temperature is abnormal, and the network port signal is abnormal.

[0073] The preset temperature range is used to indicate the temperature of a network port under normal operating conditions. Therefore, the server can determine whether the port temperature is within the preset temperature range. If the port temperature is within the preset temperature range, the network port temperature is considered normal; if the port temperature is not within the preset temperature range, the network port temperature is considered abnormal.

[0074] For example, assuming the port temperature is 30 degrees Celsius and the preset temperature range is 20 to 45 degrees Celsius, the server determines that the network port temperature is normal.

[0075] The preset parameter range is used to indicate the signal parameters of a network port when it is operating normally. Therefore, the server can determine whether the port signal parameters are within the preset parameter range. If the port signal parameters are within the preset parameter range, the network port signal is determined to be normal; if the port signal parameters are not within the preset parameter range, the network port signal is determined to be abnormal.

[0076] Assuming the port signal parameters are a signal reception power of -10dBm and a preset parameter range of greater than or equal to -50dBm, the server determines that the network port signal is normal.

[0077] In some embodiments, if the port status is normal, the network port temperature is normal, and the network port signal is normal, it indicates that the network port hardware is operating stably, the physical connection is intact, and data transmission is normal. Therefore, the server can determine that the network status is normal.

[0078] If at least one of the following conditions is present: abnormal port status, abnormal network port temperature, or abnormal network port signal, it indicates a problem affecting the normal operation of the network port storage. This could manifest as a loose physical connection of the network port, data transmission delay, or excessively high temperature. These issues may obstruct network interaction between the management board and the storage node controller, preventing normal control of the storage nodes. Therefore, the server determines the network status to be abnormal.

[0079] S303. Determine the first state based on the connection state and network state.

[0080] In some embodiments, the server determines the first state based on the connection state and the network state in the following ways: when the connection state is an abnormal connection state and / or the network state is an abnormal network state, the first state is determined to indicate that there is a fault in the network link at the current moment; when the connection state is a normal connection state and the network state is a normal network state, the first state is determined to indicate that there is no fault in the network link at the current moment.

[0081] A normal connection status indicates that the network links between the management board and the controllers of each storage node are complete and the data transmission channels are unobstructed. A normal network status indicates that the physical connection of the management board's network ports is reliable, the temperature is within a safe operating range, and the signal parameters meet stable transmission standards. Therefore, when both the connection and network status are normal, it means that the core communication requirements between the management board and the storage nodes, such as issuing control commands and transmitting status data, can be achieved smoothly. Thus, the server determines the first status to indicate that there is no network link failure at the current moment.

[0082] An abnormal connection status indicates that the network link or data transmission channel between the management board and the controllers of each storage node is blocked. An abnormal network status indicates that the management board's network port meets at least one of the following conditions: unreliable physical connection, temperature outside the safe operating range, or signal parameters not meeting stable transmission standards. Therefore, when the connection status is abnormal and / or the network status is abnormal, it means that network interaction between the management board and the storage nodes cannot proceed stably. Thus, the server determines the first status to indicate that the network link is faulty at the current moment.

[0083] exist Figure 3In the illustrated embodiment, the server first decomposes the link status information into the first connection status of each storage node controller and the second connection status of the management board. The connection status is determined through a dual verification of "all nodes are connected normally and the management board itself is connected normally." Then, the network status is determined by combining "all indicators are normal" for port status, temperature, and signal parameters. Finally, "both connection status and network status are normal" is used as the fault-free criterion, and "any abnormal status" is used as the fault criterion. This reduces the possibility of missed or incorrect fault detections caused by single-dimensional monitoring and improves the accuracy of the server in determining the first status. Furthermore, the server clarifies the basis for fault determination through the above-mentioned layered verification logic. In the event of a network link failure, the root cause can be located more accurately and quickly, making network link maintenance easier for staff.

[0084] The above embodiments describe a network management method for the management board by a server. In some embodiments, the server also stores network link status records of the management board to record its historical status. Therefore, after determining a first status, the server can update the status record information of the management board based on the first status to obtain updated status record information.

[0085] Below, in conjunction with Figure 4 The method for updating the status record information of the management board in the embodiments of this application to obtain the updated status record information will be further explained.

[0086] Figure 4 This is a schematic diagram illustrating a process for updating the status record information of a management board, provided as an embodiment of this application. Figure 4 As shown, the process may include the following steps:

[0087] S401. Obtain the status record information of the network link; the status record information includes the status of the network link at at least one historical moment.

[0088] The status record information is stored in the storage system as abnormal status information of the network link at at least one historical moment. The status record information includes the root cause of the fault at at least one historical moment and the fault repair information at at least one historical moment. For example, for any one of the at least one historical moment, the status record information may include information such as the root cause of the network link fault at that historical moment and whether the fault has been repaired.

[0089] For example, assuming that at least one historical moment includes historical moment a, historical moment b, and historical moment c, the state record information can be as shown in Table 1:

[0090] Table 1

[0091]

[0092] S402. Determine the second state of the management board based on the status record information; the second state is used to indicate whether there was a fault in the network link at the previous moment of the current moment.

[0093] In some embodiments, the server may determine whether the status record information includes the previous moment. If the status record information includes the previous moment, the server determines that the second status is used to indicate that the network link at the previous moment was faulty. If the status record information does not include the previous moment, the server determines that the second status is used to indicate that the network link at the previous moment was not faulty.

[0094] For example, assuming the status record information is as shown in Table 1, and the previous time before the current time is historical time c, the server determines the second status to indicate that there was a network link failure at the previous time.

[0095] S403. Update the state record information according to the first state and the second state to obtain the updated state record information.

[0096] In some embodiments, the server updates the status record information as follows: if a first state indicates that there is no fault in the network link at the current time and a second state indicates that there was a fault in the network link at the previous time, the fault repair information in the status record information at the previous time is updated to indicate that the fault has been repaired; if the first state indicates that there is a fault in the network link at the current time, at least one first fault root cause that caused the network link storage fault at the current time is determined, and the status record information is updated according to the at least one fault root cause to obtain the updated status record information.

[0097] Fault repair information indicates whether a network link fault at a historical point in time has been repaired. Fault repair information can indicate whether the fault has been repaired or not.

[0098] In the case where the first state indicates that there is no fault in the network link at the current moment, and the second state indicates that there was a fault in the network link at the previous moment, it means that the fault in the network link at the previous moment has been repaired. At this time, the server can update the fault repair information in the status record information of the previous moment to "fault repaired".

[0099] In some embodiments, the server determines the root cause of the fault as follows: If the first state indicates a network link failure at the current moment, and the cause of the network link failure is a port temperature exceeding a preset temperature range, the server determines the root cause to be overheating; if the cause of the network link failure is a port signal parameter below a preset parameter range, the server determines the root cause to be a low port signal; if the cause of the network link failure is a port status abnormality, the server determines the root cause to be an abnormal network port connection; if the cause of the network link failure is a connection status abnormality, the server determines the root cause to be an abnormal network link between the management board and the controller of at least one storage node.

[0100] For example, suppose the first state indicates a network link failure at the current moment, and the server determines that at least one root cause of the failure is an abnormal port state. The state record information is shown in Table 1. The server then updates the state record table based on at least one root cause of the failure, resulting in the updated state record table shown in Table 2.

[0101] Table 2

[0102]

[0103] exist Figure 4 In the illustrated embodiment, the network management method of the management board provided in this application achieves accurate location and intelligent management of network link faults by introducing a dynamic comparison mechanism between historical status records and the current status. The server automatically identifies fault repair events and updates the records to ensure the real-time nature and accuracy of status information; at the same time, when a fault is detected, it can automatically locate and record the root cause of the fault, providing clear diagnostic basis for operation and maintenance personnel and improving the efficiency and pertinence of fault handling.

[0104] In some embodiments, when the first state indicates that the network link is faulty at the current moment and the second state indicates that the network link was faulty at the previous moment, the server can also determine at least one second root cause of the fault that caused the network link to be faulty at the previous moment based on the state record information; and determine whether to generate prompt information corresponding to at least one first root cause of the fault based on at least one first root cause of the fault and at least one second root cause of the fault.

[0105] For example, assuming the status record information is as shown in Table 1, and the previous time is time c, the server determines that at least one status record information is that the port signal is low.

[0106] In some embodiments, the server may determine whether to generate prompt information corresponding to each of the at least one first fault root cause in the following manner: for each of the at least one first fault root cause, the following operations are performed: determining whether there is a fault root cause that is the same as the first fault root cause among the at least one second fault root cause; if there is no fault root cause that is the same as the first fault root cause among the at least one second fault root cause, generating prompt information corresponding to the first fault root cause; if there is a fault root cause that is the same as the first fault root cause among the at least one second fault root cause, not generating prompt information corresponding to the first fault root cause.

[0107] For example, assuming at least one first root cause of the fault includes excessively high temperature and low port signal strength, and at least one second root cause of the fault is excessively high temperature, the server determines not to generate a prompt message corresponding to excessively high temperature, but instead generates a prompt message corresponding to low port signal strength. In this case, the prompt message generated by the server is: "Low port signal strength indicates a network link failure on the management board."

[0108] In some embodiments, after determining the prompt information corresponding to the first root cause of the fault, the server can also determine the fault level of the first root cause of the fault based on the mapping relationship between the root cause of the fault and the fault level, and determine the method of sending the prompt information based on the mapping relationship between the fault level and the prompting method, as well as the fault level of the first root cause of the fault.

[0109] In some embodiments, the fault level is used to indicate the severity of the fault. The fault level can be, for example, level 1, level 2, or level 3, wherein a higher fault level indicates a greater severity of the fault.

[0110] The methods for sending notifications may include pop-up notifications, SMS notifications, email notifications, etc.

[0111] For example, suppose the mapping relationship between fault level and prompting method is as shown in Table 3:

[0112] Table 3

[0113]

[0114] If the fault level of the first root cause is level 3, the server will send a notification message via SMS.

[0115] In some embodiments, when there are many prompt messages to be sent, the sending order of the multiple prompt messages can be determined according to the fault level of the first fault root cause corresponding to each of the multiple prompt messages, and the prompt messages of the first fault root cause with higher fault level are sent first.

[0116] In the above embodiments, the server combines multi-level fault classification and differentiated alarm strategies to determine the fault level based on the root cause of the fault and send alert messages in different ways accordingly. This ensures that critical faults are handled first, while reducing interference from non-critical faults, thereby significantly improving operational efficiency, the targeted nature of fault handling, and the overall reliability and automation level of the storage system.

[0117] Figure 5 This is a schematic diagram of the structure of a network management device for a management board provided in an embodiment of this application. Figure 5 As shown, embodiments of this application also provide a network management device 50 for a management board, which includes a determination module 51, a detection module 52, and a processing module 53, wherein:

[0118] The determination module 51 is used to determine the link status information of the management board; wherein the management board is connected to the controller of at least one storage node through network links, and the link status information is used to indicate the connection status between the management board and the controller of at least one storage node.

[0119] The detection module 52 is used to detect the port status, port temperature, and port signal parameters of the network port of the management board; wherein, the network port is the port on the management board that is connected to the network;

[0120] The processing module 53 is used to determine the first state of the management board based on the link status information, port status, port temperature and port signal parameters. The first state is used to indicate whether there is a fault in the network link of the management board at the current moment.

[0121] In one possible implementation, the processing module 53 is specifically used for:

[0122] Based on the link status information, determine the connection status of the network link; the connection status is either normal connection status or abnormal connection status.

[0123] The network status of the management board is determined based on the port status, port temperature, and port signal parameters; the network status is either normal or abnormal.

[0124] The first state is determined based on the connection state and network state.

[0125] In one possible implementation, the link state information includes a first connection state of the controller of at least one storage node and a second connection state of the management board; the processing module 53 is specifically used for:

[0126] For each storage node in at least one storage node, determine whether the first connection status of the storage node's controller is normal;

[0127] Determine if the second connection is in a normal state;

[0128] The connection status is determined to be a normal connection status if the first connection status of the controllers of at least one storage node is normal and the second connection status is normal.

[0129] The connection state is determined to be an abnormal connection state if the first connection state of the controller of at least one storage node is abnormal, and / or the second connection state is abnormal.

[0130] In one possible implementation, the processing module 53 is specifically used for:

[0131] Determine whether the network port temperature is normal based on the port temperature and the preset temperature range;

[0132] Determine whether the network port signal is normal based on the port signal parameters and the preset parameter range;

[0133] If the port status is normal, the network port temperature is normal, and the network port signal is normal, then the network status is determined to be a normal network status.

[0134] The network state is determined to be abnormal if at least one of the following conditions exists: the port status is abnormal, the network port temperature is abnormal, or the network port signal is abnormal.

[0135] In one possible implementation, the processing module 53 is specifically used for:

[0136] In the case that the connection status is an abnormal connection status and / or the network status is an abnormal network status, the first status is determined to indicate that there is a fault in the network link at the current moment;

[0137] When the connection status is normal and the network status is normal, the first status is determined to indicate that there is no fault in the network link at the current moment.

[0138] In one possible implementation, the network management device 50 of the management board further includes an update module, which is specifically used for:

[0139] Obtain the status record information of the network link; the status record information includes the abnormal state of the network link at at least one historical moment;

[0140] Based on the status log information, the second status of the management board is determined; the second status is used to indicate whether there was a network link failure in the previous time.

[0141] Based on the first state and the second state, the state record information is updated to obtain the updated state record information.

[0142] In one possible implementation, the status record information includes at least one root cause of the fault at each historical moment and at least one fault repair information at each historical moment; the update module is specifically used for:

[0143] When the first state indicates that the network link is not faulty at the current time and the second state indicates that the network link was faulty at the previous time, the fault repair information at the previous time in the status record information is updated to "fault repaired";

[0144] When the first state indicates that the network link is faulty at the current time, at least one first root cause of the fault that causes the network link to be faulty at the current time is determined, and the state record information is updated according to the at least one root cause of the fault to obtain the updated state record information.

[0145] In one possible implementation, when the first state indicates a network link failure at the current moment and the second state indicates a network link failure at the previous moment, the network management device 50 of the management board further includes a prompting module, which is specifically used for:

[0146] Based on the status log information, identify at least one secondary root cause of the network link failure that occurred at the previous moment.

[0147] Based on at least one first root cause and at least one second root cause, determine whether to generate a prompt message corresponding to each of the first root causes.

[0148] In one possible implementation, the prompting module is specifically used for:

[0149] For each of the first root causes in at least one first root cause, perform the following operations:

[0150] Determine whether at least one second root cause of failure is the same as the first root cause of failure.

[0151] If there is no fault root cause that is the same as the first fault root cause among at least one second fault root cause, generate a prompt message corresponding to the first fault root cause.

[0152] If at least one of the second root causes is the same as the first root cause, no prompt message corresponding to the first root cause will be generated.

[0153] For a description of the features of the network management device 50 in the embodiment of the management board, please refer to the relevant description of the network management method in the embodiment of the management board, which will not be repeated here.

[0154] Figure 6 A schematic diagram of the structure of the electronic device provided in this application. Figure 6 As shown, the electronic device 60 provided in this embodiment includes at least one processor 601 and a memory 602. Optionally, the electronic device 60 further includes a communication component 603. The processor 601, memory 602, and communication component 603 are connected via a bus.

[0155] In the specific implementation process, at least one processor 601 executes computer execution instructions stored in memory 602, causing at least one processor 601 to execute the network management method embodiment of the management board described above.

[0156] The specific implementation process of processor 601 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0157] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0158] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0159] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0160] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above-described network management method embodiments of the management board when it is run.

[0161] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0162] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the network management method embodiments of the management board described above.

[0163] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps in any of the above-described network management method embodiments of the management board.

[0164] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0165] The network management method and electronic device for a management board provided in this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only for the purpose of helping to understand the method and its core ideas. It should be noted that those skilled in the art can make several improvements and modifications to this application without departing from the principles of this application, and these improvements and modifications also fall within the protection scope of the claims of this application.

Claims

1. A network management method for a management board, characterized in that, The method includes: Determine the link status information of the management board; wherein the management board is connected to the controller of at least one storage node via a network link, and the link status information is used to indicate the connection status between the management board and the controller of the at least one storage node; The port status, port temperature, and port signal parameters of the network port on the management board are detected; wherein, the network port is the port on the management board that is connected to the network; Based on the link status information, the port status, the port temperature, and the port signal parameters, a first state of the management board is determined. The first state is used to indicate whether there is a fault in the network link of the management board at the current moment.

2. The method according to claim 1, characterized in that, Determining the first state of the management board based on the link status information, the port status, the port temperature, and the port signal parameters includes: The connection status of the network link is determined based on the link status information; the connection status is either a normal connection status or an abnormal connection status. The network status of the management board is determined based on the port status, the port temperature, and the port signal parameters; the network status is either a normal network status or an abnormal network status. The first state is determined based on the connection state and the network state.

3. The method according to claim 2, characterized in that, The link status information includes the first connection status of the controller of each of the at least one storage node and the second connection status of the management board; determining the connection status of the network link based on the link status information includes: For each of the at least one storage node, determine whether the first connection status of the controller of the storage node is normal; Determine if the second connection status is normal; If the first connection state of the controllers of at least one storage node is normal and the second connection state is normal, the connection state is determined to be the normal connection state. If the first connection state of the controller of each of the at least one storage node is abnormal, and / or the second connection state is abnormal, the connection state is determined to be the abnormal connection state.

4. The method according to claim 2 or 3, characterized in that, Determining the network status of the management board based on the port status, the port temperature, and the port signal parameters includes: Based on the port temperature and the preset temperature range, determine whether the temperature of the network port is normal; Based on the port signal parameters and the preset parameter range, determine whether the network port signal is normal; If the port status is normal, the network port temperature is normal, and the network port signal is normal, then the network status is determined to be the normal network status. The network state is determined to be the abnormal network state if at least one of the following conditions exists: the port state is abnormal, the network port temperature is abnormal, and the network port signal is abnormal.

5. The method according to claim 2 or 3, characterized in that, Determining the first state based on the connection state and the network state includes: In the case that the connection state is the abnormal connection state and / or the network state is the abnormal network state, the first state is determined to indicate that the network link is faulty at the current time; When the connection status is the normal connection status and the network status is the normal network status, the first status is determined to indicate that there is no fault in the network link at the current time.

6. The method according to any one of claims 1-3, characterized in that, The method further includes: Obtain the status record information of the network link; the status record information is used to indicate the abnormal state of the network link at at least one historical moment; Based on the status record information, the second status of the management board is determined; the second status is used to indicate whether the network link had a fault in the time preceding the current time. Based on the first state and the second state, the state record information is updated to obtain the updated state record information.

7. The method according to claim 6, characterized in that, The status record information includes the root cause of the fault at each of the at least one historical time and the fault repair information at each of the at least one historical time; updating the status record information according to the first status and the second status to obtain the updated status record information includes: When the first state indicates that the network link is not faulty at the current time and the second state indicates that the network link was faulty at the previous time, the fault repair information at the previous time in the status record information is updated to "fault repaired"; When the first state indicates that the network link is faulty at the current time, at least one first root cause of the fault that causes the network link to be faulty at the current time is determined, and the state record information is updated according to the at least one root cause of the fault to obtain the updated state record information.

8. The method according to claim 7, characterized in that, When the first state indicates that the network link is faulty at the current time and the second state indicates that the network link was faulty at the previous time, the method further includes: Based on the status record information, at least one second root cause of failure that caused the network link to fail at the previous moment is determined. Based on the at least one first root cause of the fault and the at least one second root cause of the fault, determine whether to generate a prompt message corresponding to each of the at least one first root cause of the fault.

9. The method according to claim 8, characterized in that, The step of determining whether to generate a prompt message corresponding to each of the at least one first fault root cause and the at least one second fault root cause includes: For each of the at least one first root cause of failure, the following operations are performed: Determine whether any of the at least one second root cause of failure is the same as the first root cause of failure. If none of the at least one second fault root cause is the same as the first fault root cause, a prompt message corresponding to the first fault root cause is generated. If the same fault root cause as the first fault root cause exists among the at least one second fault root cause, no prompt information corresponding to the first fault root cause will be generated.

10. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to implement the steps of the network management method of the management board as described in any one of claims 1 to 9 when executing the computer program.