Device monitoring system and method, product, device, and storage medium

By using a monitoring scheme where the status monitor and the central processing unit serve as backups for each other, the complexity and high cost of white-box switch systems are resolved. This achieves stable and efficient monitoring when the control bus malfunctions, reduces hardware costs, and improves system reliability.

WO2026092095A1PCT designated stage Publication Date: 2026-05-07INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INSPUR SUZHOU INTELLIGENT TECH CO LTD
Filing Date
2025-10-11
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

In existing technologies, the device monitoring system of white-box switches uses two status monitors for regulation, which increases the complexity and hardware cost of the system and reduces the reliability and stability of the system.

Method used

A monitoring scheme is adopted in which the status monitor and the central processing unit serve as backups for each other. Under normal circumstances, the status monitor controls the low-speed serial bus, and in case of an anomaly, the central processing unit takes over and handles access conflicts through an arbitrator, thus avoiding the need for additional hardware.

Benefits of technology

In the event of a single control bus failure, this system prevents monitoring functions from malfunctioning, reduces hardware costs, minimizes performance loss, and improves system stability and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025127082_07052026_PF_FP_ABST
    Figure CN2025127082_07052026_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application relate to the technical field of information, and specifically relate to a device monitoring system and method, a product, a device, and a storage medium, aiming at ensuring efficient and stable operation of device monitoring systems. The system comprises: a state monitor, used for monitoring an operation state of at least one target device in a device, the target devices being connected to the state monitor; a central processing unit, used for taking over the monitoring of the target devices when the operation state of the state monitor is abnormal, the central processing unit being connected to the state monitor; and arbiters, used for controlling signals received by the target devices, the target devices being connected to the state monitor and the central processing unit by means of the arbiters.
Need to check novelty before this filing date? Find Prior Art

Description

A device monitoring system, method, product, device, and storage medium

[0001] Cross-references to related applications

[0002] This application claims priority to Chinese Patent Application No. 202411523482.9, filed on October 29, 2024, entitled “An Equipment Monitoring System, Method, Product, Equipment and Storage Medium”, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of information technology, and more specifically, to an equipment monitoring system, method, product, device, and storage medium. Background Technology

[0004] White-box switches are widely used network devices that can deploy customized applications and services according to user needs, and are particularly prevalent in large data centers. During operation, white-box switches require real-time monitoring and control of their various components. In related technologies, two status monitors are used to regulate these components; if one monitor fails, the other can be used to control the remaining components.

[0005] The method of using two status monitors to regulate the various components of a white-box switch in related technologies increases the complexity of the entire switch system structure, reduces the reliability and stability of the switch system structure, and also increases the hardware cost of the white-box switch. Summary of the Invention

[0006] This application provides an equipment monitoring system, method, product, device, and storage medium, aiming to ensure the efficient and stable operation of the equipment monitoring system.

[0007] The first aspect of this application provides a device monitoring system, the system comprising:

[0008] A status monitor is used to monitor the operating status of at least one target device in the equipment, and the target device is connected to the status monitor.

[0009] The central processing unit (CPU) is used to take over the monitoring of the target device in case of abnormal operation of the status monitor. The CPU is connected to the status monitor.

[0010] The arbiter is used to control the signals received by the target device, which is connected to the status monitor and the central processing unit through the arbiter.

[0011] In some embodiments of this application, the system further includes:

[0012] An expander is used to increase the number of interfaces of a central processing unit (CPU). The expander connects to the CPU and the arbiter.

[0013] In some embodiments of this application, the arbitrator forwards the signal sent by the status monitor to the target device when the status monitor is operating normally;

[0014] When the status monitor malfunctions, the arbiter forwards the signal sent by the central processing unit to the target device.

[0015] A second aspect of this application provides a device monitoring method, the method comprising:

[0016] At least one target device in the equipment is monitored using a status monitor;

[0017] In the event of an anomaly in any control bus of the status monitor, the central processing unit monitors the target device corresponding to the control bus.

[0018] In some embodiments of this application, the method further includes:

[0019] Start the status monitor program;

[0020] The status monitor monitoring program monitors the running status of the status monitor in real time.

[0021] In some embodiments of this application, the method further includes:

[0022] Perform read / write tests on the registers of the status monitor, and determine that the status monitor has lost its response if the registers cannot be read or written normally;

[0023] The target device corresponding to the status monitor is monitored through the control bus of the central processing unit.

[0024] In some embodiments of this application, the method further includes:

[0025] Once the status monitor's operating status is detected to have returned to normal, resume the status monitor's monitoring of the target device.

[0026] In some embodiments of this application, the recovery status monitor's monitoring of the target device includes:

[0027] Shut down the control bus corresponding to the central processing unit.

[0028] Disable the program interface corresponding to the central processing unit;

[0029] Delete the information corresponding to the control bus that caused the error from the database;

[0030] The target device is monitored using the control bus corresponding to the status monitor.

[0031] In some embodiments of this application, monitoring at least one target device in the device is performed using a status monitor, including:

[0032] The operating status data of each target device is read through the status monitor;

[0033] The status monitor sends the operating status data to the central processing unit.

[0034] In some embodiments of this application, in the event of an anomaly in any control bus of the status monitor, the central processing unit monitors the target device corresponding to the control bus, including:

[0035] Determine the control bus of the status monitor corresponding to the target device;

[0036] Turn off the control bus corresponding to the status monitor;

[0037] The target device is monitored using the control bus corresponding to the central processing unit.

[0038] In some embodiments of this application, the target device is monitored using the control bus corresponding to the central processing unit, including:

[0039] Load the driver program corresponding to the control bus of the central processing unit;

[0040] The signal sent by the central processing unit is forwarded to the target device by the arbiter corresponding to the target device.

[0041] In some embodiments of this application, the method further includes:

[0042] The status monitor program sends out corresponding alarm information.

[0043] Record any abnormal information about the control bus corresponding to the target device in the log.

[0044] In some embodiments of this application, the method further includes:

[0045] The number of control buses that malfunctioned and their control bus numbers are written to the database to inform programs that depend on the control buses to switch the program interface of the status monitor to the program interface of the central processing unit.

[0046] A third aspect of this application provides a device monitoring apparatus, the apparatus comprising:

[0047] The first monitoring module is used to monitor at least one target device in the equipment through a status monitor;

[0048] The first monitoring takeover module is used to monitor the target device corresponding to any control bus in the event of an anomaly in any control bus of the status monitor through the central processing unit.

[0049] In some embodiments of this application, the apparatus further includes:

[0050] The program startup module is used to start the status monitor program;

[0051] The status monitor module is used to monitor the running status of the status monitor in real time through the status monitor monitoring program.

[0052] In some embodiments of this application, the apparatus further includes:

[0053] The monitoring recovery module is used to restore the monitoring of the target device by the status monitor when the status monitor's operating status is detected to have returned to normal.

[0054] In some embodiments of this application, the first monitoring module includes:

[0055] The status data reading submodule is used to read the operating status data of each target device through the status monitor;

[0056] The data sending submodule is used to send operational status data to the central processing unit via the status monitor.

[0057] In some embodiments of this application, the monitoring takeover module includes:

[0058] The bus determination submodule is used to determine the control bus of the status monitor corresponding to the target device;

[0059] The bus shutdown submodule is used to shut down the control bus corresponding to the status monitor;

[0060] The monitoring and control submodule is used to monitor the target device using the control bus corresponding to the central processing unit.

[0061] In some embodiments of this application, the monitoring takeover submodule includes:

[0062] The driver loading submodule is used to load the driver program corresponding to the control bus of the central processing unit.

[0063] The signal forwarding submodule is used to forward the signals sent by the central processing unit to the target device through the arbiter corresponding to the target device.

[0064] In some embodiments of this application, the apparatus further includes:

[0065] The alarm module is used to monitor the program through the status monitor to issue corresponding alarm information;

[0066] The exception information recording module is used to record exception information of the control bus corresponding to the target device in the log.

[0067] In some embodiments of this application, the apparatus further includes:

[0068] The data writing module is used to write the number of control buses that have malfunctioned and their control bus numbers into the database, so as to inform programs that depend on the control buses to switch the program interface of the status monitor to the program interface of the central processing unit.

[0069] In some embodiments of this application, the apparatus further includes:

[0070] The read / write test module is used to perform read / write tests on the registers of the status monitor and determine that the status monitor has lost its response if the registers cannot be read or written normally.

[0071] The central processing unit (CPU) takes over the monitoring module and is used to monitor the target devices corresponding to the status monitor through the CPU's control bus.

[0072] In some embodiments of this application, the monitoring and recovery module includes:

[0073] The bus shutdown submodule is used to shut down the control bus corresponding to the central processing unit;

[0074] The program interface shutdown submodule is used to shut down the program interface corresponding to the central processing unit;

[0075] The information deletion submodule is used to delete information corresponding to the control bus that has malfunctioned from the database;

[0076] The device monitoring submodule is used to monitor the target device using the control bus corresponding to the status monitor.

[0077] A fourth aspect of this application provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the method of the first aspect of this application.

[0078] The fifth aspect of this application provides a non-volatile readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method as described in the first aspect of this application.

[0079] A sixth aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the method of the first aspect of this application.

[0080] The equipment monitoring system provided in this application includes a status monitor for controlling the operating status of at least one target device of the equipment, each target device being connected to the status monitor via a serial bus; a central processing unit (CPU) for taking over control of the target device corresponding to the control bus when the operating status of the status monitor is abnormal, the CPU being connected to the status monitor; an arbitrator for controlling the signals received by the target device, the target device being connected to the status monitor and the CPU via the arbitrator; and an expander for expanding the number of interfaces of the CPU, the expander being connected to the CPU and the arbitrator.

[0081] In this equipment monitoring system, the operating status of at least one target device in the switch is controlled first through the status monitor. If any control bus in the status monitor fails and cannot control one or more target devices, the central processing unit takes over the corresponding target device. This prevents the white-box switch monitoring function from becoming unavailable due to status monitor malfunction. The white-box switch low-speed bus controller, which uses the central processing unit as a backup, does not require additional hardware components, reducing the hardware cost of the white-box switch. The failure of a single control bus will not cause the overall function to be migrated, reducing the performance loss when a single control bus malfunctions. Attached Figure Description

[0082] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0083] Figure 1 is a structural diagram of an equipment monitoring system proposed in an embodiment of this application;

[0084] Figure 2 is a flowchart of a device monitoring method according to an embodiment of this application;

[0085] Figure 3 is a schematic diagram of the operation flow of a monitoring program according to an embodiment of this application;

[0086] Figure 4 is a schematic diagram of a device monitoring device according to an embodiment of this application;

[0087] Figure 5 is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0088] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0089] Referring to Figure 1, which is a structural diagram of an equipment monitoring system according to an embodiment of this application, the system includes:

[0090] A status monitor controls the operating status of at least one target device in the switch. Each target device is connected to the status monitor via a serial bus. A central processing unit (CPU) takes over control of the target device corresponding to the control bus in case of an abnormal status in the status monitor. The CPU is connected to the status monitor. An arbitrator controls the signals received by the target devices. The target devices are connected to the status monitor and the CPU via the arbitrator. An expander increases the number of interfaces on the CPU. The expander is connected to the CPU and the arbitrator. If the number of target devices to be monitored is small and the CPU has enough interfaces, the expander is not required. When the status monitor is operating normally, the arbitrator forwards the signals sent by the status monitor to the target devices. When the status monitor is operating abnormally, the arbitrator forwards the signals sent by the CPU to the target devices.

[0091] In this embodiment, each control bus includes an arbiter and each control bus corresponds to a target device. For example, arbiter 1 corresponds to target device 1, arbiter 2 corresponds to target device 2, and arbiter 3 corresponds to target device 3. The number of control buses is the same as the number of target devices that need to be monitored. The number of target devices that need to be monitored will determine the number of control buses between the target device and the status monitor, as well as the number of control buses between the target device and the central processing unit.

[0092] In this embodiment, the device can be a server, a white-box switch, or similar equipment. This example will be illustrated using a white-box switch. A white-box switch is an open network device with decoupled hardware and software, characterized by flexibility, efficiency, and programmability, and can significantly reduce network deployment costs. A white-box switch consists of two parts: hardware and software. The hardware generally includes switching processing components, a CPU (Central Processing Unit), a network interface card (NIC), storage, and peripheral hardware devices. The software refers to the NOS (Network Operating System) and its network applications. A status monitor is a CPLD (Complex Programmable Logic Device) used to monitor the operating status of at least one target device in a switch. The status monitor connects to the devices in the switch via I2C (Inter-Integrated Circuit) and to the central processing unit (CPU) via PCIe (Peripheral Component Interconnect Express). It sends data read from the target device to the CPU. An arbiter controls the signals received by the target device; each control bus has one arbiter. To prevent access conflicts between the CPU and the status monitor, the arbiter forwards signals. When the status monitor is operating normally, signals sent by the status monitor are preferentially forwarded to the target device. Target devices are typically various sensors, fans, power supplies, etc., and the temperature of various hardware components can be monitored through these sensors. An expander is used to expand the CPU's interface. One end of the expander connects to the CPU, and the other end connects to the arbiter, which then connects to the target device.

[0093] For example, the central processing unit (CPU) uses an ARM (Advanced RISC Machine) architecture. A CPLD (Portable Logic Controller) acts as a status monitor, retrieving information from devices on the low-speed serial bus (I2C) instead of the CPU. The status monitor communicates with the CPU via the PCIe bus. After obtaining device information through its own I2C controller, the status monitor caches it in its own memory. Programs on the operating system retrieve information through the high-speed PCIe serial bus. The CPU's I2C controller is also connected to devices on each low-speed serial bus. If the number of channels needs to be expanded, a switch can be used as an expander. Each I2C bus is connected to both the status monitor and the CPU's I2C controller. Each bus has an I2C arbitrator to handle potential access conflicts and system crashes.

[0094] In this embodiment, a monitoring scheme in which the status monitor and the central processing unit (CPU) act as backups for each other is adopted. That is, all low-speed devices of the white-box switch need to be connected to the low-speed bus controller of the CPU and the status monitor at the same time. Under normal circumstances, the status monitor is used to control the low-speed serial bus. When any control bus in the status monitor malfunctions, the control of that control bus will be transferred to the CPU under software control. When the status monitor malfunctions, the control of all control buses corresponding to the status monitor will be transferred to the CPU to ensure that the hardware control does not fail in case of malfunction. The backup function of the low-speed serial bus device control of the white-box switch is realized without adding additional devices.

[0095] In this embodiment, the hardware devices that the white-box switch needs to monitor, in addition to being connected to the low-speed serial bus of the status monitor, also need to be connected to the low-speed serial bus of the central processing unit. Under normal circumstances, the low-speed serial bus of the status monitor is used first, and the low-speed serial bus controller of the central processing unit is turned off. In order to prevent possible access conflicts between the two parties and to handle the situation where the line is suspended due to a fault of one party, the devices on the low-speed bus are first connected to the arbiter, and then connected to the status monitor and the central processing unit.

[0096] Referring to Figure 2, which is a flowchart of a device monitoring method according to an embodiment of this application, the method includes the following steps:

[0097] S11: Monitor at least one target device in the equipment using a status monitor.

[0098] In this embodiment, when the switch is started, under normal circumstances, at least one target device in the switch is first monitored by the status monitor. During the monitoring process, the status monitor sends a data read signal to the arbitrator. The status monitor sends the data read signal to the target device, and the target device then returns the data to the status monitor. The status monitor sends the received data to the central processing unit. After receiving the data from the target device, the central processing unit performs the corresponding adjustment.

[0099] In this embodiment, the step of monitoring at least one target device in the switch through a status monitor specifically includes:

[0100] S11-1: Read the operating status data of each target device through the status monitor.

[0101] In this embodiment, when the switch starts up, the operating status data of each target device is read by the status monitor. When reading the data, the data reading request is forwarded by the arbiter corresponding to the target device. After receiving the data reading request, the target device sends the corresponding data to the status monitor.

[0102] For example, operational status data includes device temperatures detected by various temperature sensors, fan speed, power supply power, etc.

[0103] S11-2: The status monitor sends the operating status data to the central processing unit.

[0104] In this embodiment, the status monitor and the central processing unit are connected via a high-speed serial bus. After reading the operating status data of the target device, the status monitor sends the operating status data to the central processing unit.

[0105] S12: In the event of an anomaly in any control bus of the status monitor, the target device corresponding to the control bus is monitored through the central processing unit.

[0106] In this embodiment, when one or more low-speed serial buses in the status monitor fail and cannot transmit information normally, or when a certain interface of the status monitor fails, the status monitor monitoring program will shut down the faulty control bus and then use the central processing unit to take over the monitoring of the corresponding target device. When the central processing unit takes over the monitoring, the driver corresponding to the central processing unit is loaded in the target device, and then the program interface is switched to the program interface corresponding to the central processing unit.

[0107] In this embodiment, the specific steps for monitoring the target device through the central processing unit include:

[0108] S12-1: Determine the control bus of the status monitor corresponding to the target device.

[0109] In this embodiment, each target device is controlled by a control bus, which is a low-speed serial bus. When a control bus fails, the control bus corresponding to the monitored target device must first be determined.

[0110] In this embodiment, each target device has a corresponding control bus number, and each target device also has a corresponding device number. The control bus number of each control bus is pre-recorded in the system. When the data of a target device cannot be read, the control bus that is currently malfunctioning can be determined based on the pre-recorded control bus number of each control bus.

[0111] S12-2: Turn off the control bus corresponding to the status monitor.

[0112] In this embodiment, after identifying the faulty control bus, the control bus on the status monitor used to read the target device is turned off.

[0113] S12-3: Monitor the target device using the control bus corresponding to the central processing unit.

[0114] In this embodiment, after the control bus of the status monitor is turned off, the control bus corresponding to the central processing unit is used to monitor the target device.

[0115] In this embodiment, the specific steps for monitoring the target device using the control bus corresponding to the central processing unit include:

[0116] S12-3-1: Load the driver program corresponding to the control bus of the central processing unit.

[0117] In this embodiment, the bus controller of the central processing unit and the bus controller of the status monitor have different drivers. When the central processing unit takes over the monitoring of a target device, the monitoring program running in the background loads the driver of the control bus corresponding to the central processing unit. After the driver is loaded, the devices in the switch can transmit data with the central processing unit.

[0118] S12-3-2: The signal sent by the central processing unit is forwarded to the target device through the arbiter corresponding to the target device.

[0119] In this embodiment, when the arbiter receives a signal from the central processing unit, and the control bus corresponding to the status monitor fails, the arbiter sends the signal sent by the central processing unit to the corresponding target device, thus completing the central processing unit's monitoring of the target device.

[0120] In this embodiment, when any control bus of the status monitor malfunctions, the target device corresponding to the malfunctioning control bus is handed over to the central processing unit for monitoring. The other control buses in the status monitor are unaffected, and there is no need to migrate all the target devices monitored by the status monitor as a whole. This makes more reasonable resource allocation and saves the central processing unit's resources.

[0121] In this embodiment, the method further includes:

[0122] S12-4: The status monitor program issues corresponding alarm information.

[0123] In this embodiment, when the status monitor program detects an abnormality in the status monitor, it issues a corresponding alarm message, which includes information such as the number of fault control buses and the fault control bus number.

[0124] S12-5: Record the abnormal information of the control bus corresponding to the target device in the log.

[0125] In this embodiment, when an anomaly is detected in the status monitor, the abnormal information of the control bus corresponding to the target device is recorded in the log, the control bus number of the control bus corresponding to the target device is recorded, and the number of control buses that have failed is recorded.

[0126] In this embodiment, when the status monitor detects an anomaly, it promptly issues an alarm and records the anomaly information in the log. This allows administrators to obtain the anomaly information in a timely manner and to check the log to determine the bus where the anomaly occurred, which is beneficial for rapid system maintenance.

[0127] In this embodiment, the method further includes:

[0128] S12-6: Write the number of control buses that have malfunctioned and their control bus numbers into the database to inform programs that depend on the control buses to switch the program interface of the status monitor to the program interface of the central processing unit.

[0129] In this embodiment, when the control bus of the status monitor, i.e. the low-speed serial bus, malfunctions, the number of control buses that malfunction and their control bus numbers are written into the database. After the other programs that depend on the control bus obtain the information about the malfunction of the control bus from the database, they switch to the backup code and switch the program interface of the status monitor to the program interface of the central processing unit.

[0130] In this embodiment, when any related device or other program in the switch reads data from the target device, the program interface is switched from the status monitor's program interface to the central processing unit's program interface, and the target device is accessed through the central processing unit's program interface.

[0131] In this embodiment, in addition to the monitoring program of the status monitor, other programs running on the white-box switch operating system that need to control the target device corresponding to the low-speed bus also need to be backed up. Since the way to access the target device through the low-speed serial bus of the central processing unit is different from that through the low-speed serial bus of the status monitor, the driving paths are different. Therefore, after these programs discover the abnormal status of the low-speed serial bus reported by the status monitor program through various means, such as database or sysfs node, they will automatically switch program interfaces and use the backed-up code to access the target device corresponding to the low-speed serial bus.

[0132] In this embodiment, timely sharing of control bus information that has encountered an anomaly through the database facilitates timely switching of corresponding program interfaces by other programs that depend on the target device, thereby improving the system's working efficiency.

[0133] In another embodiment of this application, when the switch is powered on, the method further includes:

[0134] S21: Start the status monitor program.

[0135] In this embodiment, the status monitor monitoring program is used to monitor the running status of the status monitor.

[0136] In this embodiment, when the switch is started, the status monitor begins to monitor at least one target device in the switch. At the same time, the system background starts the status monitor monitoring program to monitor the status monitor.

[0137] In this embodiment, the network operating system of the white-box switch continuously checks the running status of the status monitor through the status monitor monitoring program. When an anomaly occurs in the low-speed serial bus controlled by the status monitor, the monitoring program immediately alarms and records the anomaly log, and retains the operating environment data at the time of the anomaly for easy use in problem analysis. At the same time, if the anomaly of the status monitor causes anomalies such as stuttering on the low-speed serial bus, the arbiter on the low-speed serial bus will automatically send a signal to try to restore the transmission of the low-speed serial bus. If the anomaly cannot be restored, in order to ensure that the low-speed serial bus controller of the central processing unit takes over the monitoring task, the monitoring program will also load the driver program of the device managed by the low-speed serial bus controller of the central processing unit and create the corresponding program interface for other devices that depend on the program to use the interface.

[0138] S22: The status monitor monitoring program monitors the running status of the status monitor in real time.

[0139] In this embodiment, after the status monitor monitoring program is started, the running status of the status monitor is monitored in real time through the status monitor monitoring program.

[0140] In this embodiment, the operational status of the status monitor is monitored in two main ways. One method involves reading and writing to the status monitor's registers. If a register cannot be read or written, the status monitor is considered unresponsive, indicating a failure of the entire status monitor and requiring a complete migration of monitoring. The other method records the status of the low-speed serial bus for each status monitor. If the target device corresponding to the low-speed serial bus remains unresponsive for a certain period, or if it continuously reports errors for an extended period, the low-speed serial bus is considered unresponsive. In this case, the central processing unit (CPU) takes over the monitoring of the target device on the faulty control bus.

[0141] In this embodiment, the status monitor is monitored in real time by the status monitor monitoring program, which helps to keep track of the status monitor's operating status at all times. When the status monitor malfunctions, corresponding measures can be taken in a timely manner, thereby improving the system's operating efficiency.

[0142] In another embodiment of this application, the method further includes:

[0143] S31: If the status monitor is found to have returned to normal operation, resume the status monitor's monitoring of the target device.

[0144] In this embodiment, when the status monitor program detects that the status monitor's operating status has returned to normal, the status monitor resumes monitoring of the target device.

[0145] In this embodiment, the status monitor program runs continuously. While the central processing unit takes over the monitoring of the target device, the status monitor program continuously sends signals to the status monitor through the arbiter of the fault control bus to automatically attempt to restore the bus's monitoring of the target device. When the status monitor responds, it indicates that the status monitor has returned to normal operation, and at this time, the status monitor's monitoring of the target device is restored.

[0146] In this embodiment, the recovery status monitor's monitoring of the target device includes:

[0147] S31-1: Shut down the control bus corresponding to the central processing unit.

[0148] In this embodiment, when the control bus of the status monitor is detected to have returned to normal, the monitoring task needs to be migrated back to the status monitor to ensure that the operation of the central processing unit is not affected. At this time, the control bus corresponding to the central processing unit is shut down first.

[0149] S31-2: Disable the program interface corresponding to the central processing unit.

[0150] In this embodiment, while shutting down the control bus corresponding to the central processing unit, the program interface corresponding to the central processing unit in the target device is also shut down. At this time, the target device can no longer interact with the central processing unit.

[0151] S31-3: Delete the information corresponding to the control bus that caused the exception from the database.

[0152] In this embodiment, when a control bus failure occurs, the database records information such as the control bus number of the failed control bus. When the system returns to normal, the information corresponding to the abnormal control bus is deleted from the database.

[0153] S31-4: Use the control bus corresponding to the status monitor to monitor the target device.

[0154] In this embodiment, the control bus corresponding to the status monitor is used to monitor the target device. At this time, the monitoring task of the target device is transferred from the central processing unit to the status monitor.

[0155] In this embodiment, when the status monitor program detects that the status monitor has returned to normal operation, or when the control bus in the status monitor that previously failed has returned to normal, the monitoring task of the target device is transferred from the central processing unit to the status monitor, thus ensuring the operating speed of the central processing unit.

[0156] In another embodiment of this application, the method further includes:

[0157] S41: Perform read / write tests on the registers of the status monitor, and determine that the status monitor has lost its response if the registers cannot be read or written normally.

[0158] In this embodiment, in order to test whether the status monitor is operating normally, the register of the status monitor is read and written when the switch is started. Data is repeatedly written to and read from the register of the status monitor by the arbiter of each low-speed serial bus control bus. If the data can be read and written normally, it indicates that the status monitor is operating normally.

[0159] S42: Monitors the target device corresponding to the status monitor through the control bus of the central processing unit.

[0160] In this embodiment, if the registers in the status monitor cannot be read or written normally, it is determined that the status monitor has lost its response.

[0161] In this embodiment, when it is determined that the status monitor has lost its response, all control buses in the status monitor are unavailable. Therefore, all target devices monitored by the status monitor are handed over to the central processing unit for monitoring. All control buses of the status monitor are shut down, and the driver of the low-speed serial bus controller of the central processing unit is loaded into each target device and the remaining programs. The program interface is switched, and the monitoring task is fully migrated to the central processing unit.

[0162] In this embodiment, when a fault is detected in the status monitor, all target devices corresponding to the status monitor are promptly handed over to the central processing unit for monitoring, ensuring the stable operation of the system.

[0163] Referring to Figure 3, which is a schematic diagram of the monitoring program operation flow according to an embodiment of this application, after the white-box switch's operating system starts, the status monitor program first checks whether the status monitor can be read and written normally. If it cannot be read and written normally, the backup driver is directly loaded, an alarm is triggered, the fault information is recorded in the log, and the fault information is written to the database to inform other programs. If it can be read and written normally, it determines whether there is an abnormal low-speed bus in the status monitor. If it exists, the abnormal status monitor path is closed, the backup controller driver is loaded to open the backup bus, an alarm is triggered, the abnormal information is recorded in the log, and other programs that depend on the bus are notified. The number of control buses and the control bus number of the abnormal control bus are written to the operating system's database. After the low-speed serial bus is abnormal, the monitoring program will continue to monitor the status of the status monitor. If it finds that the abnormal function of the status monitor has recovered and continues for a period of time, it attempts to remove the abnormal state and restore the original low-speed bus controller and corresponding code. The duration of the abnormal function recovery can be set according to the actual situation and is not limited here.

[0164] In another embodiment of this application, when a control bus of the status monitor malfunctions and the central processing unit takes over the corresponding target device, the arbiter attempts to restore communication of the control bus within a preset time period. If communication is not restored within the preset time period, the system records the failure of the control bus restoration in the log and marks the control bus with a maintenance mark. When the staff checks the log, they can know that the control bus cannot be restored automatically. At this time, the control bus is repaired. After the repair, the maintenance mark in the log is removed. The system detects the removal of the maintenance mark and restores the control bus again until the control bus returns to normal operation.

[0165] In the above embodiments of this application, the status monitor and the central processing unit (CPU) serve as backups for each other, monitoring each device in the white-box switch. This prevents the white-box device monitoring function from failing due to anomalies in the status monitor. Furthermore, using the CPU as a backup low-speed bus controller for the white-box switch eliminates the need for additional hardware, reducing hardware design complexity and saving costs. When a single low-speed bus malfunctions, it does not lead to the migration of the entire function; only the monitoring task of a single target device needs to be migrated, improving system efficiency, reducing performance loss caused by anomalies, and promptly handling and automatically attempting to restore the abnormal control bus. Once the abnormal control bus returns to normal, it can be reactivated, ensuring the efficiency of the CPU and thus optimizing the overall efficiency of the white-box switch operating system.

[0166] Based on the same inventive concept, one embodiment of this application provides an equipment monitoring device. Referring to FIG4, FIG4 is a schematic diagram of an equipment monitoring device 400 according to an embodiment of this application. As shown in FIG4, the device includes:

[0167] The first monitoring module 401 is used to monitor at least one target device in the equipment through a status monitor;

[0168] The first monitoring takeover module 402 is used to monitor the target device corresponding to the control bus through the central processing unit in the event of any abnormality in the control bus of the status monitor.

[0169] In some embodiments of this application, the apparatus further includes:

[0170] The program startup module is used to start the status monitor program;

[0171] The status monitor module is used to monitor the running status of the status monitor in real time through the status monitor monitoring program.

[0172] In some embodiments of this application, the apparatus further includes:

[0173] The monitoring recovery module is used to restore the monitoring of the target device by the status monitor when the status monitor's operating status is detected to have returned to normal.

[0174] In some embodiments of this application, the first monitoring module includes:

[0175] The status data reading submodule is used to read the operating status data of each target device through the status monitor;

[0176] The data sending submodule is used to send operational status data to the central processing unit via the status monitor.

[0177] In some embodiments of this application, the monitoring takeover module includes:

[0178] The bus determination submodule is used to determine the control bus of the status monitor corresponding to the target device;

[0179] The bus shutdown submodule is used to shut down the control bus corresponding to the status monitor;

[0180] The monitoring and control submodule is used to monitor the target device using the control bus corresponding to the central processing unit.

[0181] In some embodiments of this application, the monitoring takeover submodule includes:

[0182] The driver loading submodule is used to load the driver program corresponding to the control bus of the central processing unit.

[0183] The signal forwarding submodule is used to forward the signals sent by the central processing unit to the target device through the arbiter corresponding to the target device.

[0184] In some embodiments of this application, the apparatus further includes:

[0185] The alarm module is used to monitor the program through the status monitor to issue corresponding alarm information;

[0186] The exception information recording module is used to record exception information of the control bus corresponding to the target device in the log.

[0187] In some embodiments of this application, the apparatus further includes:

[0188] The data writing module is used to write the number of control buses that have malfunctioned and their control bus numbers into the database, so as to inform the programs that depend on the control buses to switch the program interface of the status monitor to the program interface of the central processing unit.

[0189] In some embodiments of this application, the apparatus further includes:

[0190] The read / write test module is used to perform read / write tests on the registers of the status monitor and determine that the status monitor has lost its response if the registers cannot be read or written normally.

[0191] The central processing unit (CPU) takes over the control module, which is used to monitor the target device corresponding to the status monitor through the CPU's control bus.

[0192] In some embodiments of this application, the monitoring and recovery module includes:

[0193] The bus shutdown submodule is used to shut down the control bus corresponding to the central processing unit;

[0194] The program interface shutdown submodule is used to shut down the program interface corresponding to the central processing unit;

[0195] The information deletion submodule is used to delete information corresponding to the control bus that has malfunctioned from the database;

[0196] The device monitoring submodule is used to monitor the target device using the control bus corresponding to the status monitor.

[0197] Based on the same inventive concept, another embodiment of this application provides a non-volatile readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the device monitoring method of any of the above embodiments of this application.

[0198] Based on the same inventive concept, another embodiment of this application provides an electronic device. FIG5 is a schematic diagram of an electronic device 500 proposed in an embodiment of this application, including a memory 502, a processor 501 and a computer program stored in the memory and executable on the processor. When executed by the processor, the steps in the device monitoring method of any of the above embodiments of this application are implemented.

[0199] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0200] Some embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to mutually.

[0201] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, embodiments of this application can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of this application can take the form of a computer program product implemented on one or at least one computer-usable storage medium (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0202] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one block of the flowchart illustration or at least one block of the block diagram or at least one block.

[0203] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0204] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in a flowchart, one or more flowcharts, and / or a block diagram, one or more blocks.

[0205] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0206] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes the element.

[0207] The above provides a detailed description of the equipment monitoring system, method, product, equipment, and storage medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A device monitoring system, characterized in that, The system includes: A status monitor is configured to monitor the operating status of at least one target device in the equipment, the target device being connected to the status monitor; A central processing unit (CPU) is configured to take over monitoring the target device in the event of an abnormal operating state of the status monitor; the CPU is connected to the status monitor. The arbiter is configured to control the signals received by the target device, which is connected to the status monitor and the central processing unit via the arbiter.

2. The equipment monitoring system according to claim 1, characterized in that, The system also includes: An expander is configured to increase the number of interfaces of the central processing unit, and the expander is connected to the central processing unit and the arbiter.

3. The equipment monitoring system according to claim 1, characterized in that, When the status monitor is operating normally, the arbitrator forwards the signal sent by the status monitor to the target device. When the status monitor is in an abnormal operating state, the arbitrator forwards the signal sent by the central processing unit to the target device.

4. A method for monitoring equipment, characterized in that, The method includes: At least one target device in the device is monitored by a status monitor; In the event of an anomaly in any control bus of the status monitor, the target device corresponding to the control bus is monitored by the central processing unit.

5. The equipment monitoring method according to claim 4, characterized in that, After the step of monitoring the target device corresponding to any control bus of the status monitor in the event of an anomaly, the method further includes: The communication of the control bus is restored within a preset time period by the arbitrator; If the communication of the control bus is not restored by the arbiter within the preset time period, the failure to restore the communication of the control bus is recorded in the log, and the control bus is marked with a maintenance mark.

6. The equipment monitoring method according to claim 4, characterized in that, The method further includes: Start the status monitor program; The status monitor monitoring program monitors the running status of the status monitor in real time.

7. The equipment monitoring method according to claim 6, characterized in that, The method further includes: The registers of the status monitor are read and written, and if the registers cannot be read or written normally, it is determined that the status monitor has lost its response. The target device corresponding to the status monitor is monitored through the control bus of the central processing unit.

8. The equipment monitoring method according to claim 6, characterized in that, The method further includes: The status of the low-speed serial bus of the status monitor is obtained, and if the target device corresponding to the low-speed serial bus does not respond within a preset time period or the time for continuously reporting errors exceeds a preset time, it is determined that the low-speed serial bus has lost response. The central processing unit monitors the target device corresponding to the low-speed serial bus.

9. The equipment monitoring method according to claim 4, characterized in that, The method further includes: If the status monitor is found to have returned to normal operation, the status monitor shall resume monitoring of the target device.

10. The equipment monitoring method according to claim 9, characterized in that, Restoring the monitoring of the target device by the status monitor includes: Shut down the control bus corresponding to the central processing unit; Close the program interface corresponding to the central processing unit; Delete the information corresponding to the control bus of the status monitor that caused the anomaly from the database; The target device is monitored using the control bus corresponding to the status monitor.

11. The equipment monitoring method according to claim 4, characterized in that, The monitoring of at least one target device in the device via a status monitor includes: The operating status data of the target device is read through the status monitor; The status monitor sends the operating status data to the central processing unit.

12. The equipment monitoring method according to claim 11, characterized in that, The operating status data includes device temperature detected by the temperature sensor, fan rotation speed, and power supply power.

13. The equipment monitoring method according to claim 4, characterized in that, In the event of an anomaly on any control bus of the status monitor, the central processing unit monitors the target device corresponding to the control bus, including: Determine the control bus of the status monitor corresponding to the target device; Turn off the control bus of the corresponding status monitor; The target device is monitored using the control bus corresponding to the central processing unit.

14. The equipment monitoring method according to claim 13, characterized in that, The step of monitoring the target device using the control bus corresponding to the central processing unit includes: Load the driver program corresponding to the control bus of the central processing unit; The signal sent by the central processing unit is forwarded to the target device by the arbiter corresponding to the target device.

15. The equipment monitoring method according to claim 13, characterized in that, The method further includes: The status monitor program sends out corresponding alarm information. Record the abnormal information of the status monitor control bus corresponding to the target device in the log.

16. The equipment monitoring method according to claim 15, characterized in that, The alarm information includes the number of fault control buses and the number of the fault control buses.

17. The equipment monitoring method according to claim 13, characterized in that, The method further includes: The number of abnormal state monitor control buses and their control bus numbers are written to the database to inform programs that depend on the state monitor control buses to switch the state monitor's program interface to the central processing unit's program interface.

18. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 4-17.

19. A computer non-volatile readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 4 to 17.

20. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 4 to 17.

Citation Information

Patent Citations

  • Switch monitoring device and method based on programmable logic device

    CN112019455A

  • Baseboard management controller fault detection device

    CN114691408A

  • BMC upgrading method and device

    CN116028094A

  • Switch operation control method, device, system and equipment and storage medium

    CN116319618A

  • Server exception monitoring method and device, equipment and storage medium

    CN116846790A