Part power-on and power-off state monitoring method, device and equipment and storage medium

By capturing and saving the up-down power state information of components in the storage system, and using the dynamic rule engine to perform abnormal detection, the shortcomings of up-down power state monitoring of components in the storage system are solved, accurate recording and rapid fault positioning of component states are achieved, and operation and maintenance efficiency is improved.

CN120492247APending Publication Date: 2025-08-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510603117.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The prior art cannot fully monitor the power up and down state of components in the storage system cluster, making it difficult to trace the root cause of the failure and prone to false alarms or missed alarms.

Method used

The up and down state information of the component is captured through the hardware interface and software driver, and saved to the component power status register, and abnormal detection is performed using the dynamic rule engine.

Benefits of technology

Real-time monitoring and accurate event recording of the power up and down states of components in the storage system is realized, reducing the frequency of manual inspection, quickly positioning abnormal components, and extending the overall life of the storage system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120492247A_ABST
    Figure CN120492247A_ABST
Patent Text Reader

Abstract

The invention discloses a part power-on and power-off state monitoring method, device and equipment and a storage medium, and relates to the technical field of storage, and the method comprises the steps: capturing the power-on and power-off state information of parts in each slot through a hardware interface and a software driver, achieving the real-time full-amount obtaining of the power-on and power-off state information of all slot parts, and improving the reliability of the part power-on and power-off state information. Accurate event recording of the power-on and power-off state of each component is realized by storing the obtained power-on and power-off state information of each component to the preset component power state register. The dynamic rule engine is utilized to automatically perform component anomaly detection, so that the manual inspection frequency is reduced, and the abnormal component can be quickly positioned. The technical problems that the power-on and power-off states of the components in the system cluster cannot be comprehensively mastered, the fault source is difficult to trace, and false alarm or missing alarm is caused are solved, and the technical effects that the power-on and power-off state information of all the slot position components can be fully obtained, accurate event recording of the power-on and power-off states of all the components is achieved, and abnormal components can be rapidly positioned are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of storage technology, and in particular to a method, device, equipment, and storage medium for monitoring the power-on and power-off status of components. Background Art

[0002] Storage systems are the core infrastructure of data centers, and their stability directly impacts business continuity. Storage systems contain numerous pluggable components, such as Fibre Channel (FC) cards, Serial Attached SCSI (SAS) cards, and hard drives. Abnormal power cycles on these components can lead to data loss, hardware damage, or system downtime.

[0003] Currently, the commonly used method for monitoring the power-on and power-off status of components is to monitor the power-on and power-off status of a single component. This method cannot fully grasp the power-on and power-off status of components in the system cluster, cannot predict possible fault conditions, and makes it difficult for operation and maintenance personnel to trace the root cause of the fault, resulting in false alarms or missed alarms. Summary of the Invention

[0004] The present application provides a component power-on and power-off status monitoring method that can fully obtain the power-on and power-off status information of all slot components, realize accurate event recording of the power-on and power-off status of each component, and quickly locate abnormal components, so as to at least solve the problem in related technologies that the power-on and power-off status of components in the system cluster cannot be fully grasped, the root cause of the fault is difficult to trace, and false alarms or missed alarms may occur.

[0005] This application provides a method for monitoring the power-on and power-off status of a component, including:

[0006] Capturing power-up and power-down status information of components in each slot through a hardware interface and software driver; wherein the power-up and power-down status information includes a timestamp when the power-up and power-down status of the component changes;

[0007] Saving the power-on and power-off status information of each component to a preset component power status register, so as to collect statistics on the power-on and power-off status information corresponding to each component using the component power status register;

[0008] The dynamic rule engine is used to retrieve the power-on and power-off status information of each component obtained by statistics from the component power status register to perform component abnormality detection.

[0009] The present application also provides a component power-on and power-off status monitoring device, comprising:

[0010] A power-up and power-down status information acquisition module, configured to capture the power-up and power-down status information of components in each slot through a hardware interface and software driver; wherein the power-up and power-down status information includes a timestamp when the power-up and power-down status of the component changes;

[0011] A power-up and power-down status information statistics module is used to save the power-up and power-down status information of each component to a preset component power status register, so as to use the component power status register to collect statistics on the power-up and power-down status information corresponding to each component;

[0012] The component abnormality detection module is used to use a dynamic rule engine to retrieve the power-on and power-off status information of each component obtained by statistics from the component power status register to perform component abnormality detection.

[0013] The present application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for implementing the steps of any of the above-mentioned component power-on and power-off status monitoring methods when executing the computer program.

[0014] The present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned component power-on and power-off status monitoring methods are implemented.

[0015] Through this application, since the power-on and power-off status information of the components in each slot can be captured through the hardware interface and software driver, real-time monitoring of the power-on and power-off status of the components in the storage system can be achieved, and the power-on and power-off status information of all slot components can be fully obtained. By saving the obtained power-on and power-off status information of each component to a preset component power status register, the power-on and power-off status information of the component is statistically analyzed, and accurate event records of the power-on and power-off status of each component are achieved. By automatically detecting component anomalies using a dynamic rule engine, the frequency of manual inspections is reduced, abnormal components can be quickly located, the fault location time is greatly reduced, labor costs are saved, and the operational efficiency of component maintenance is improved. Furthermore, aging components can be replaced in advance based on the component anomaly detection results to extend the overall life of the storage system. Therefore, it is possible to solve the technical problem of not being able to fully grasp the power-on and power-off status of components in the system cluster, and it is difficult to trace the root cause of the fault, resulting in false alarms or missed reports, and achieve the technical effect of being able to fully obtain the power-on and power-off status information of all slot components, achieving accurate event records of the power-on and power-off status of each component, and quickly locating abnormal components. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A flowchart of a method for monitoring the power-on and power-off status of a component provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of another component power-on and power-off status monitoring method provided in an embodiment of the present application;

[0019] Figure 3 A timing diagram of a component power-on and power-off status monitoring process provided in an embodiment of the present application;

[0020] Figure 4 This is a structural block diagram of a component power-on and power-off status monitoring device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0022] It should be noted that, in the description of this application, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. The terms "first," "second," etc., in this application are used to distinguish similar objects, and are not used to describe a particular order or sequence.

[0023] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0024] In conjunction with the specific application environment architecture or specific hardware architecture on which the execution of the component power-on and power-off status monitoring method depends, the specific application environment architecture or specific hardware architecture is described herein.

[0025] An embodiment of the present application provides a component power-on and power-off status monitoring method, and the method is described in detail in conjunction with the execution flow of the component power-on and power-off status monitoring method.

[0026] See also Figure 1 , Figure 1 This is a flowchart of a method for monitoring the power-on and power-off status of a component provided in an embodiment of the present application. The method may include the following steps:

[0027] S101: Capture power-on and power-off status information of components in each slot through a hardware interface and software driver.

[0028] The power-on and power-off status information includes a timestamp when the power-on and power-off status of a component changes.

[0029] Pre-configured hardware interfaces and software drivers are used to capture power-on and power-off events during system operation. These interfaces can include the Intelligent Platform Management Interface (IPMI) and General-Purpose Input / Output (GPIO) interfaces, and software drivers can include Userspace Device Management Rules (udev) and kernel modules. During system operation, these interfaces capture the power-on and power-off status of components in each slot. This captured power-on and power-off status information includes the timestamp of the component power-on and power-off status change.

[0030] The components may include FC cards, SAS cards, network cards, hard disks, memory, power supply units (PSUs), baseband units (BBUs), and so on.

[0031] S102: Saving the power-on and power-off status information of each component to a preset component power status register, so as to collect statistics on the power-on and power-off status information corresponding to each component using the component power status register.

[0032] A component power status register (FRU_power) is pre-set to store the power-up and power-down status information of each component. After capturing the power-up and power-down status information of the components in each slot, the power-up and power-down status information of each component is saved to the pre-set component power status register. The component power status register is then used to collect statistics on the power-up and power-down status information corresponding to each component.

[0033] S103: Utilize a dynamic rule engine to retrieve the power-on and power-off status information of each component obtained through statistics from the component power status register to perform component abnormality detection.

[0034] A dynamic rule engine is pre-configured to detect component anomalies based on the power-up and power-down status information of each component. When a component anomaly detection trigger condition is met, such as when a component anomaly detection request is received or when a preset component anomaly detection cycle is reached, the dynamic rule engine retrieves the statistical power-up and power-down status information of each component from the component power status register to perform component anomaly detection.

[0035] Through this application, since the power-on and power-off status information of the components in each slot can be captured through the hardware interface and software driver, real-time monitoring of the power-on and power-off status of the components in the storage system can be achieved, and the power-on and power-off status information of all slot components can be fully obtained. By saving the obtained power-on and power-off status information of each component to a preset component power status register, the power-on and power-off status information of the component is statistically analyzed, and accurate event records of the power-on and power-off status of each component are achieved. By automatically detecting component anomalies using a dynamic rule engine, the frequency of manual inspections is reduced, abnormal components can be quickly located, the fault location time is greatly reduced, labor costs are saved, and the operational efficiency of component maintenance is improved. Furthermore, aging components can be replaced in advance based on the component anomaly detection results to extend the overall life of the storage system. Therefore, it is possible to solve the technical problem of not being able to fully grasp the power-on and power-off status of components in the system cluster, and it is difficult to trace the root cause of the fault, resulting in false alarms or missed reports, and achieve the technical effect of being able to fully obtain the power-on and power-off status information of all slot components, achieving accurate event records of the power-on and power-off status of each component, and quickly locating abnormal components.

[0036] See also Figure 2 , Figure 2 This is a flowchart of another method for monitoring the power-on and power-off status of a component provided in an embodiment of the present application. The method may include the following steps:

[0037] S201: responding to hardware signals of various components in real time through an interrupt service routine pre-registered in an error correction code module of the storage system.

[0038] Pre-register an interrupt service routine (ISR) in the storage system's error correction code (EC) module. This pre-registered ISR responds to hardware signals from various components in real time.

[0039] S202: At the storage system hardware layer, hardware signals of the power-on and power-off status of each component are read through the slot power status register, the presence detection pin, the input voltage and output current registers to obtain power-on and power-off status information of the components in each slot.

[0040] The power-on and power-off status information includes a timestamp when the power-on and power-off status of a component changes.

[0041] At the storage system hardware layer, hardware signals indicating the power-up and power-down status of each component are read through the slot power status register, presence detection pin, and input voltage and output current registers. This information is then captured for each component in each slot. This information is captured by configuring the storage system error correction code module and the storage system hardware layer. This provides effective software and hardware support for capturing component power-up and power-down status information, enabling comprehensive acquisition of power-up and power-down status information for each component.

[0042] The slot power status register can be set to the Peripheral Component Interconnect Express (PCIe) slot power status register (Offset 0x3E). The input voltage and output current registers can be set to the input voltage (VIN) and output current (IOUT) registers of the Power Management Bus (PMBus) protocol.

[0043] S203: The power-on and power-off status information of each component obtained from the band is saved in the component power status register, so as to perform array analysis on the power-on and power-off status information of each component obtained by statistics to obtain a binary array list.

[0044] After reading the power-on and power-off status information of each component from the band, the power-on and power-off status information of each component obtained from the band is saved in the component power status register, and then the power-on and power-off status information of each component obtained by statistics is analyzed in the component power status register to obtain a binary array list.

[0045] S204: Utilize the component power status register to store the binary array list.

[0046] After the binary array list is statistically generated, it is stored in the component power status register. By analyzing the power-on and power-off status information of each component, the binary array list is stored. For example, the binary array list can be generated by time dimension (daily, weekly, or monthly) to assist in operation and maintenance decision-making. Component power-on and power-off status information can be filtered, viewed, and managed according to actual needs, achieving effective statistical analysis of each component's power-on and power-off status information and improving storage space utilization.

[0047] Since the power-on and power-off status information of each component obtained in-band and out-of-band and stored in the component power status register will be overwritten over time, it can be set to store the power-on and power-off status information of each component obtained in-band and out-of-band in a log file, thereby avoiding data loss and facilitating subsequent data search.

[0048] S205: Retrieve a binary array list from the component power status register and send it to the webpage graphical user interface.

[0049] After storing the binary array list in the component power status register, when the component power-on / off status information detection trigger condition is met, the binary array list is retrieved from the component power status register and sent to the web graphical user interface (GUI WEB). By setting up the web GUI, user interaction with the system is achieved, and users can manage system commands and data through the graphical interface.

[0050] S206: Mapping the binary array list into a device status matrix of power-on and power-off status information of each component on the webpage graphical user interface.

[0051] After the binary array list is sent to the web GUI, it is mapped into a device status matrix showing the power-on and power-off status of each component. The background color of the device status matrix cards can be dynamically adjusted based on the component's power-on and power-off status (normal / abnormal), providing a visual display of each component's power-on and power-off status. A scrolling refresh of the latest 10 event records can also be configured to be pulled from the database every second. Clicking on a detail will redirect to the device history page, allowing users to view historical power-on and power-off status information for each component at any time.

[0052] S207: Split the power-on and power-off status information of the components in the device state matrix, and receive the updated power-on and power-off status information of the components in real time according to the split result to update the device state matrix.

[0053] After the webpage graphical user interface maps the binary array list into a device status matrix containing the power-up and power-down status information for each component, the component power-up and power-down status information in the device status matrix is split. This allows the device status matrix to be updated in real time based on the split results. By splitting the power-up and power-down status information for each component in the device status matrix, the display of each component's power-up and power-down status information can be controlled independently, improving the flexibility of controlling the webpage graphical user interface in displaying each component's power-up and power-down status information.

[0054] Regarding the statistical analysis view in the web graphical user interface, the embodiment of the present application adopts a time series database and a rolling window algorithm to achieve efficient data storage and multi-dimensional statistics. It can be linked to charts and displayed as categories such as FC / SAS, network card, PSU, BBU, hard disk, etc. After the user selects the time range (day / week / month), the front end sends an AJAX request to obtain data, and the back end returns the aggregated results. The front end dynamically renders the chart through ECharts. Daily statistics aggregate data once every hour to generate a 24-hour rolling report. Weekly / monthly statistics are based on the rolling window (Tumbling Window) algorithm and are displayed in fixed periodic summary data. Interactive charts support zooming, dragging and drill-down queries, and access data export. Click the "Export CSV" button to call the back end to generate a temporary file and return a download link.

[0055] S208: The storage system error correction code module calls the baseboard management controller command to obtain the power-on and power-off status information of each component from out-of-band, and saves the power-on and power-off status information of each component obtained from out-of-band to the component power status register.

[0056] In addition to obtaining the power-on and power-off status information of each component in-band, the power-on and power-off status information of each component can also be obtained out-of-band by calling the baseboard management controller command through the storage system error correction code module, that is, the power-on and power-off status information of each component is obtained through the external network, and the power-on and power-off status information of each component obtained out-of-band is saved in the component power status register.

[0057] S209: performing data deduplication, data verification, and time stamp calibration on the power-on and power-off status information corresponding to each component obtained from the in-band and the power-on and power-off status information obtained from the out-band.

[0058] After storing the power-up and power-down status information of each component obtained out-of-band in the component power status register, data deduplication, data verification, and timestamp alignment are performed on the corresponding in-band and out-of-band power-up and power-down status information for each component. This deduplication, data verification, and timestamp alignment ensures the accuracy of the power-up and power-down status information stored in the component power status register, avoids the storage of redundant data, and improves resource utilization.

[0059] In addition, a Message Digest Algorithm (MD5) hash value can be generated for the acquired data to ensure data security. Duplicate events can be filtered through log records, and the Network Time Protocol (NTP) protocol can be used to calibrate the time of each node, for example, ensuring that the error is less than 1 millisecond, thereby further ensuring data accuracy from multiple aspects.

[0060] In a specific implementation of the present application, step S209 may include the following steps:

[0061] Step 1: Obtain the timestamps of the power-on and power-off status changes of the components contained in the power-on and power-off status information obtained in-band and out-of-band, respectively.

[0062] Step 2: Get the system time corresponding to the power-up and power-down status information of each component obtained from the in-band and the power-up and power-down status information obtained from the out-band;

[0063] Step 3: Use each system time to calibrate the timestamps of the power-on and power-off states of the corresponding components to obtain the timestamp calibration results;

[0064] Step 4: Determine, based on the timestamp calibration result, whether there is power-up and down status information obtained in-band and out-of-band corresponding to the same component obtained at the same time point. If so, proceed to step 5; if not, do not remove the power-up and down status information.

[0065] Step 5: Select the power status information to be removed from the two power status information according to the consistency determination result of the two power status information, and remove the power status information to be removed from the component power status register.

[0066] For the convenience of description, the above five steps can be combined for explanation.

[0067] During data deduplication, data verification, and timestamp calibration, the timestamps of component power-up and down status changes, contained in the in-band and out-of-band power-up and down status information corresponding to each component, can be obtained. The system time corresponding to each in-band and out-of-band power-up and down status information can be obtained, and the timestamps of the corresponding component power-up and down status changes can be calibrated using each system time to obtain a timestamp calibration result. Based on the timestamp calibration result, it is determined whether the in-band and out-of-band power-up and down status information corresponding to the same component are obtained at the same time. If so, step five is executed. If not, the power-up and down status information is not removed. Based on the consistency determination result of the two power-up and down status information, the power-up and down status information to be removed is selected from the two power-up and down status information, and the power-up and down status information to be removed is removed from the component power status register. By calibrating the timestamps of component power-up and down status changes using the system time, the accuracy of the timestamps of the power-up and down status changes of each component is ensured. When determining that both in-band and out-of-band power status information corresponding to the same component exist at the same time, the system performs a consistency check on the in-band and out-of-band power status information. Based on the consistency check result, the system selects the power status information to be removed from the two pieces of information and removes the removed power status information from the component power status register. This ensures accurate screening of valid power status information and timely removal of redundant power status information.

[0068] In a specific embodiment of the present application, selecting the power status information to be removed from the two power status information based on the consistency determination result of the two power status information, and removing the power status information to be removed from the component power status register may include the following steps:

[0069] Step 1: When the two pieces of power-on and power-off status information are consistent, any one of the two pieces of power-on and power-off status information is removed from the component power status register;

[0070] Step 2: When the two power status information are inconsistent, the power status information obtained from out-of-band is determined as the power status information of the component at the corresponding time point, and the power status information obtained from in-band is removed from the component power status register.

[0071] For the convenience of description, the above two steps can be combined for explanation.

[0072] When the two power-up and power-down status information are consistent, any one of the two power-up and power-down status information is removed from the component power-up and power-down status register. When the two power-up and power-down status information are inconsistent, the power-up and power-down status information obtained from out-of-band is determined to be the power-up and power-down status information of the component at the corresponding time point, and the power-up and power-down status information obtained from in-band is removed from the component power-up and power-down status register. By determining that the power-up and power-down status information obtained from in-band is inconsistent with the power-up and power-down status information obtained from out-of-band, the power-up and power-down status information obtained from out-of-band is determined to be the power-up and power-down status information of the component at the corresponding time point, and the power-up and power-down status information obtained from in-band is removed from the component power-up and power-down status register. This can avoid interference from business logic and network environment, and ensure the accuracy of the component power-up and power-down status information stored in the component power-up and power-down status register.

[0073] S210: Utilizing a dynamic rule engine to retrieve statistically obtained power-on and power-off status information of each component from a component power status register to perform component abnormality detection.

[0074] When using the dynamic rule engine to detect component anomalies based on the power-on and power-off status information of each component, dynamic alarms can be triggered based on preset thresholds (such as more than three power-on and power-off times within 24 hours), improving system reliability.

[0075] After using the dynamic rule engine to retrieve the statistically obtained power-on and power-off status information of each component from the component power status register to perform component abnormality detection, the method may further include the following steps:

[0076] Step 1: Obtain component anomaly detection results for each component;

[0077] Step 2: Generate an alarm list based on the abnormality detection results of each component;

[0078] Step 3: Get the preset alarm priority;

[0079] Step 4: Output the alarm information in the alarm list according to the alarm priority.

[0080] For the convenience of description, the above four steps can be combined for explanation.

[0081] Unhandled alarms can be displayed in an alarm list, with component categories filtered by priority (high, medium, or low). Pre-configured batch operations allow you to acknowledge or ignore multiple alarms with a single click. By reporting alarms based on abnormal detection results, you can predict component lifespan and failures, extending the overall lifespan of the storage system and improving storage survivability.

[0082] In addition, you can also set the exception handling process:

[0083] (1) If there is a component that has been inserted but not added to the system in the current system, this slot will display 0x00, triggering the process of adding components. If the process of adding components exceeds 120s and still has not been added to the system, the client prompts that it has timed out, and this slot will display 0xff. The GUI WEB list page will display the information that the power-on and power-off status of the component cannot be identified;

[0084] (2) Abnormal data filtering: For illegal timestamps, events earlier than the system startup time or in the future will be discarded. For state jump detection, if the device reports the same event multiple times within 10 seconds, it will be marked as abnormal data and entered into the abnormal database and not included in the statistical scope;

[0085] (3) Automatic repair mechanism, for redundant data comparison, if the information of the same component is reported from multiple sources (such as IPMI and driver layer), the one with the earliest timestamp is taken as the valid event;

[0086] (4) Changes in power-on and power-off status: If the same component is powered on and off more than three times within 24 hours, the problem alarm status will be recorded and the alarm queue will be recorded.

[0087] See also Figure 3 , Figure 3 A timing diagram of a component power-on and power-off status monitoring process provided in an embodiment of the present application. A request for monitoring component power-on and power-off status is sent to the storage system on the GUI WEB side of the storage system. The storage system calls the Programmable Logic (PL) module interface to query the power-on and power-off status information of all components and feeds back the current power-on and power-off information of each component to the storage system. The storage system caches the component power-on and power-off status information in the component power-on and power-off status cache (Lcache). The WEB page WebSocket pushes the information data based on the current component power-on and power-off status Lcache, pulls the latest events from the database every second, achieves zero-delay refresh, and displays it as a component power-on and power-off status information view, including a real-time monitoring panel, a statistical analysis view, and an alarm management interface. The component power-on and power-off status Lcache always retains the component power-on and power-off status information. When the component power-on and power-off status information changes, the PL module reports it to the storage system in real time, and the component power-on and power-off status Lcache updates the cache, achieving real-time monitoring of the upper-layer module. When the GUI WEB requests next time, the component power-on and power-off status Lcache change information will be returned to the WEB page. The power-on and power-off status information of the components of the PL module interface is obtained through the power-on and power-off status monitoring system of the underlying EC module components.

[0088] As shown in Table 1, Table 1 is a monitoring flow chart of some components monitored by the error correction code module of the storage system.

[0089] Table 1

[0090]

[0091] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0092] An embodiment of the present application further provides a device for monitoring the power-on and power-off status of a component, which may include:

[0093] The power-up and power-down status information acquisition module 41 is used to capture the power-up and power-down status information of the components in each slot through the hardware interface and software driver; wherein the power-up and power-down status information includes the timestamp when the power-up and power-down status of the components changes;

[0094] The power-up and power-down status information statistics module 42 is used to save the power-up and power-down status information of each component to a preset component power status register, so as to use the component power status register to collect statistics on the power-up and power-down status information corresponding to each component;

[0095] The component abnormality detection module 43 is used to use a dynamic rule engine to retrieve the power-on and power-off status information of each component obtained by statistics from the component power status register to perform component abnormality detection.

[0096] Through this application, since the power-on and power-off status information of the components in each slot can be captured through the hardware interface and software driver, real-time monitoring of the power-on and power-off status of the components in the storage system can be achieved, and the power-on and power-off status information of all slot components can be fully obtained. By saving the obtained power-on and power-off status information of each component to a preset component power status register, the power-on and power-off status information of the component is statistically analyzed, and accurate event records of the power-on and power-off status of each component are achieved. By automatically detecting component anomalies using a dynamic rule engine, the frequency of manual inspections is reduced, abnormal components can be quickly located, the fault location time is greatly reduced, labor costs are saved, and the operational efficiency of component maintenance is improved. Furthermore, aging components can be replaced in advance based on the component anomaly detection results to extend the overall life of the storage system. Therefore, it is possible to solve the technical problem of not being able to fully grasp the power-on and power-off status of components in the system cluster, and it is difficult to trace the root cause of the fault, resulting in false alarms or missed reports, and achieve the technical effect of being able to fully obtain the power-on and power-off status information of all slot components, achieving accurate event records of the power-on and power-off status of each component, and quickly locating abnormal components.

[0097] In a specific embodiment of the present application, the power-on and power-off status information acquisition module 42 may include:

[0098] The hardware signal response submodule is used to respond to the hardware signals of each component in real time through the interrupt service routine pre-registered in the error correction code module of the storage system;

[0099] The power-on and power-off status information acquisition submodule is used to read the hardware signals of the power-on and power-off status of each component through the slot power status register, in-place detection pin, input voltage and output current registers at the storage system hardware layer, and obtain the power-on and power-off status information of the components in each slot.

[0100] In a specific embodiment of the present application, the power-up and power-down status information statistics module 42 is specifically a module that saves the power-up and power-down status information of each component obtained from the band to the component power status register;

[0101] The device may also include:

[0102] The power-on and power-off status information storage module is used to call the baseboard management controller command through the storage system error correction code module to obtain the power-on and power-off status information of each component from out-of-band, and save the power-on and power-off status information of each component obtained from out-of-band to the component power status register;

[0103] The data processing module is used to perform data deduplication, data verification and time stamp calibration on the power-on and power-off status information corresponding to each component obtained from the in-band and the power-on and power-off status information obtained from the out-band.

[0104] In a specific embodiment of the present application, the data processing module may include:

[0105] The timestamp acquisition submodule is used to obtain the timestamp of the power-on and power-off status change of each component contained in the power-on and power-off status information obtained from the in-band and out-of-band, respectively.

[0106] The system time acquisition submodule is used to obtain the system time corresponding to the power-on and power-off status information obtained from the in-band and out-of-band of each component;

[0107] The timestamp calibration result obtaining submodule is used to calibrate the timestamps when the power-on and power-off states of the corresponding components change using the system time to obtain the timestamp calibration results;

[0108] A judgment submodule, configured to judge, based on the timestamp calibration result, whether there is power-on and power-off status information obtained in-band and power-on and power-off status information obtained out-of-band corresponding to the same component obtained at the same time point;

[0109] The power-on and power-off status information elimination submodule is used to select the power-on and power-off status information to be eliminated from the two power-on and power-off status information according to the consistency judgment result of the two power-on and power-off status information when it is determined according to the timestamp calibration result that there is power-on and power-off status information obtained from the in-band and the power-on and power-off status information obtained from the out-band corresponding to the same component obtained at the same time point, and eliminate the power-on and power-off status information to be eliminated from the component power status register.

[0110] In a specific embodiment of the present application, the power-on and power-off status information elimination submodule may include:

[0111] A first power-up and power-down status information removing unit is configured to remove any one of the two power-up and power-down status information from the component power status register when the two power-up and power-down status information are consistent;

[0112] The second power on and off status information elimination unit is used to determine the power on and off status information obtained from out-of-band as the power on and off status information of the component at the corresponding time point when the two power on and off status information are inconsistent, and to eliminate the power on and off status information obtained from in-band from the component power status register.

[0113] In a specific embodiment of the present application, the power-on and power-off status information statistics module 42 may include:

[0114] A binary array list obtaining submodule is used to perform array analysis on the power-on and power-off status information of each component obtained by statistics to obtain a binary array list;

[0115] The binary array list storage submodule is used to store the binary array list using the component power status register.

[0116] In a specific embodiment of the present application, the device may further include:

[0117] A binary array list sending module is used to retrieve a binary array list from a component power status register and send it to a webpage graphical user interface;

[0118] A device status matrix mapping module is used to map the binary array list into a device status matrix of the power-on and power-off status information of each component in the web graphical user interface;

[0119] The device state matrix update module is used to split the power-on and power-off status information of the components in the device state matrix, and to receive the updated power-on and power-off status information of the components in real time according to the split result to update the device state matrix.

[0120] For the description of the features in the embodiment corresponding to the component power-on and power-off status monitoring device, reference can be made to the relevant description of the embodiment corresponding to the component power-on and power-off status monitoring method, which will not be repeated here.

[0121] An embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above-mentioned component power-on and power-off status monitoring method embodiments.

[0122] An embodiment of the present application further provides a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps of any of the above-mentioned component power-on and power-off status monitoring method embodiments when running.

[0123] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media that can store computer programs, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk, or an optical disk.

[0124] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the steps in any of the above-mentioned component power-on and power-off status monitoring methods are implemented.

[0125] An embodiment of the present application also provides another computer program product, including a non-volatile computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the steps in any of the above-mentioned component power-on and power-off status monitoring methods.

[0126] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0127] The above is a detailed introduction to a component power-on and power-off status monitoring method, device, equipment and storage medium provided by the present application. This article uses specific examples to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the present application.

Claims

1. A method for monitoring the power-on and power-off status of a component, characterized in that: include: Capturing power-up and power-down status information of components in each slot through a hardware interface and software driver; wherein the power-up and power-down status information includes a timestamp when the power-up and power-down status of the component changes; Saving the power-on and power-off status information of each component to a preset component power status register, so as to collect statistics on the power-on and power-off status information corresponding to each component using the component power status register; The dynamic rule engine is used to retrieve the power-on and power-off status information of each component obtained by statistics from the component power status register to perform component abnormality detection.

2. The component power-on and power-off status monitoring method according to claim 1, characterized in that: Capture power-on and power-off status information of components in each slot through hardware interfaces and software drivers, including: Respond to hardware signals of each component in real time through the interrupt service routine pre-registered in the error correction code module of the storage system; At the storage system hardware layer, the hardware signals of the power-on and power-off status of each component are read through the slot power status register, the presence detection pin, the input voltage and output current registers, and the power-on and power-off status information of the components in each slot is obtained.

3. The component power-on and power-off status monitoring method according to claim 1, characterized in that: The power-on and power-off status information of each component is saved to the preset component power status register, including: Saving the power-on and power-off status information of each component obtained from the band to the component power status register; Correspondingly, it also includes: The storage system error correction code module calls the baseboard management controller command to obtain the power-on and power-off status information of each component from out-of-band, and saves the power-on and power-off status information of each component obtained from out-of-band to the component power status register; The power-on and power-off status information corresponding to each component obtained from the in-band and the power-on and power-off status information obtained from the out-band are deduplicated, verified, and timestamp calibrated.

4. The component power-on and power-off status monitoring method according to claim 3, characterized in that: Perform data deduplication, data verification, and timestamp calibration on the power-on and power-off status information obtained in-band and out-of-band for each component, including: Obtaining the timestamps of the power-on and power-off status changes of the components contained in the power-on and power-off status information obtained in-band and out-of-band, respectively, corresponding to each component; Obtain the system time corresponding to the power-on and power-off status information of each component obtained from in-band and from out-of-band; Use each system time to calibrate the timestamps of the power-on and power-off states of the corresponding components to obtain the timestamp calibration results; Determining, based on the timestamp calibration result, whether there is power-on and power-off status information obtained in-band and power-on and power-off status information obtained out-of-band corresponding to the same component obtained at the same time point; If so, the power status information to be removed is selected from the two power status information according to the consistency determination result of the two power status information, and the power status information to be removed is removed from the component power status register.

5. The component power-on and power-off status monitoring method according to claim 4, characterized in that: Selecting the to-be-eliminated up / down state information from the two up / down state information according to a consistency determination result of the two up / down state information, and removing the to-be-eliminated up / down state information from the component power state register, including: When the two pieces of power-on and power-off status information are consistent, removing any one of the two pieces of power-on and power-off status information from the component power supply status register; When the two power status information are inconsistent, the power status information obtained from outside the band is determined as the power status information of the component at the corresponding time point, and the power status information obtained from within the band is removed from the component power status register.

6. The component power-on and power-off status monitoring method according to any one of claims 1 to 5, characterized in that: The component power status register is used to collect statistics on the power-on and power-off status information corresponding to each component, including: Perform array analysis on the power-on and power-off status information of each component obtained by statistics to obtain a binary array list; The binary array list is stored using the component power status register.

7. The component power-on and power-off status monitoring method according to claim 6, characterized in that: After storing the binary array list using the component power status register, the method further includes: Retrieving the binary array list from the component power status register and sending it to a webpage graphical user interface; Mapping the binary array list into a device status matrix of power-on and power-off status information of each component in the webpage graphical user interface; The power-on and power-off status information of the components in the device state matrix is split, and the updated power-on and power-off status information of the components is received in real time according to the split result to update the device state matrix.

8. A device for monitoring the power-on and power-off status of a component, characterized in that: include: A power-up and power-down status information acquisition module, configured to capture the power-up and power-down status information of components in each slot through a hardware interface and software driver; wherein the power-up and power-down status information includes a timestamp when the power-up and power-down status of the component changes; A power-up and power-down status information statistics module is used to save the power-up and power-down status information of each component to a preset component power status register, so as to use the component power status register to collect statistics on the power-up and power-down status information corresponding to each component; The component abnormality detection module is used to use a dynamic rule engine to retrieve the power-on and power-off status information of each component obtained by statistics from the component power status register to perform component abnormality detection.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor is configured to implement the steps of the component power-on and power-off status monitoring method according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the component power-on and power-off status monitoring method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Switch system

    CN121585631A