Edge device monitoring method, system and device and electronic device

By deploying out-of-band information acquisition modules and out-of-band monitoring methods in the edge computing environment, the problem of difficult monitoring of hardware failures of edge devices is solved, real-time perception and anomaly warning of edge devices are realized, and the reliability and operation and maintenance efficiency of the system are improved.

CN121887670APending Publication Date: 2026-04-17BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING BAIDU NETCOM SCI & TECH CO LTD
Filing Date
2025-12-17
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In edge computing scenarios, hardware failures of a large number of edge devices, which are geographically distributed and operate in complex environments, are difficult to monitor in a timely manner, affecting system stability.

Method used

By deploying an out-of-band information acquisition module, hardware health information of edge devices is obtained through the out-of-band management interface, and the data center performs out-of-band monitoring to achieve real-time perception and anomaly warning of edge devices.

Benefits of technology

It improves the visibility and operational transparency of edge devices, reduces the complexity of fault location and diagnosis, enhances the maintainability and management efficiency of devices, and ensures the high availability and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121887670A_ABST
    Figure CN121887670A_ABST
Patent Text Reader

Abstract

The invention discloses an edge device monitoring method, system and device and electronic equipment, and relates to the technical field of computers, in particular to the field of artificial intelligence such as edge computing and cloud computing. According to the specific implementation scheme, the method comprises the following steps: receiving hardware health information sent by an out-of-band information acquisition module deployed in an edge cluster; wherein the hardware health information is acquired by the out-of-band information acquisition module by performing information acquisition on hardware of edge equipment in an edge cluster through an out-of-band management interface; and performing out-of-band monitoring on the edge device according to the hardware health information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to the fields of artificial intelligence such as edge computing and cloud computing, specifically to an edge device monitoring method, system, device and electronic device. Background Technology

[0002] In edge computing scenarios, there are a vast number of edge devices, geographically distributed, and operating in complex environments. These edge devices are typically deployed in confined environments such as outdoor base stations, pole-mounted cabinets, and factory production lines. During operation, their hardware may experience hardware failures that affect system stability, such as CPU overheating, fan failure, and power supply anomalies. Therefore, in order to promptly address potential hardware risks and sudden failures, it is essential to monitor edge devices. Summary of the Invention

[0003] This application provides a method, system, apparatus, and electronic device for monitoring edge devices. The specific solution is as follows: According to one aspect of this application, an edge device monitoring method is provided, comprising: Receive hardware health information sent by the out-of-band information acquisition module deployed in the edge cluster; wherein, the hardware health information is obtained by the out-of-band information acquisition module through the out-of-band management interface by collecting information from each hardware of the edge devices in the edge cluster; Out-of-band monitoring of edge devices is performed based on hardware health information.

[0004] According to another aspect of this application, an edge device monitoring method is provided, comprising: The out-of-band information acquisition module deployed in the edge cluster collects information from each hardware device in the edge cluster through the out-of-band management interface to obtain hardware health information; Send hardware health information to the data center so that the data center can perform out-of-band monitoring of edge devices based on the hardware health information.

[0005] According to another aspect of this application, an edge device monitoring system is provided, comprising: a data center and an edge cluster deployed with out-of-band information acquisition modules; The data center is used to execute the method described in one aspect of the above embodiment, and the out-of-band information acquisition module is used to execute the method described in another aspect of the embodiment.

[0006] According to another aspect of this application, an edge device monitoring device is provided, comprising: The receiving module is used to receive hardware health information sent by the out-of-band information acquisition module deployed in the edge cluster; the hardware health information is obtained by the out-of-band information acquisition module through the out-of-band management interface by collecting information from each hardware of the edge devices in the edge cluster. The monitoring module is used to perform out-of-band monitoring of edge devices based on hardware health information.

[0007] According to another aspect of this application, an edge device monitoring device is provided, comprising: The out-of-band information acquisition module is used to collect information from the hardware of each edge device in the edge cluster through the out-of-band management interface and obtain hardware health information; The sending module is used to send hardware health information to the data center so that the data center can perform out-of-band monitoring of edge devices based on the hardware health information.

[0008] According to another aspect of this application, an electronic device is provided, comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method described in the above embodiments.

[0009] According to another aspect of this application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method described in the above embodiments.

[0010] According to another aspect of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method described in the above embodiments.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of this application. Wherein: Figure 1 A flowchart illustrating an edge device monitoring method provided in an embodiment of this application; Figure 2 A flowchart illustrating an edge device monitoring method provided in another embodiment of this application; Figure 3A flowchart illustrating an edge device monitoring method provided in another embodiment of this application; Figure 4 A flowchart illustrating an edge device monitoring method provided in another embodiment of this application; Figure 5 A flowchart illustrating an edge device monitoring method provided in another embodiment of this application; Figure 6 This is a schematic diagram of the structure of an edge device monitoring system provided in an embodiment of this application; Figure 7 A schematic diagram illustrating the interaction between an out-of-band information acquisition module of multiple edge clusters and a data center, provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an edge device monitoring device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of an edge device monitoring device provided in another embodiment of this application; Figure 10 This is a block diagram of an electronic device used to implement the edge device monitoring method of the embodiments of this application. Detailed Implementation

[0013] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] It should be noted that the acquisition, storage, use, and processing of data in this application comply with the relevant provisions of national laws and regulations and do not violate public order and good morals.

[0015] The edge device monitoring method, system, apparatus, electronic device, and storage medium according to embodiments of this application are described below with reference to the accompanying drawings.

[0016] Figure 1 This is a flowchart illustrating an edge device monitoring method provided in an embodiment of this application.

[0017] The edge device monitoring method of this application embodiment can be executed by the edge device monitoring device of this application embodiment, which can be configured in an electronic device.

[0018] Among them, electronic devices can be any device with computing capabilities, such as personal computers, mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0019] For example, the edge device monitoring method of this application embodiment can be executed by a data center, which can be used to perform out-of-band monitoring of edge devices. For example, the data center can also be described as a central management platform, and this application does not limit the name of the data center.

[0020] like Figure 1 As shown, the edge device monitoring method includes: Step 101: Receive hardware health information sent by the out-of-band information acquisition module deployed on the edge cluster.

[0021] In this application, the data center can monitor one or more edge clusters, each edge cluster contains multiple edge devices, and each edge cluster is equipped with an out-of-band information acquisition module. The data center can receive hardware health information sent by the out-of-band information acquisition modules of each edge cluster.

[0022] As can be seen, by communicating with an out-of-band information acquisition module, the data center can obtain the hardware health information of all edge devices in an edge cluster.

[0023] For example, hardware health information can be collected by an out-of-band information acquisition module, which collects information from each piece of hardware in the edge devices of the edge cluster through an out-of-band management interface. It is evident that hardware health information is obtained through a dedicated management channel—the out-of-band management interface—independent of the host operating system and the main business network, and does not rely on the main network link between the edge devices and the data center. This acquisition method offers high reliability.

[0024] For example, hardware health information can be used to describe the health status of the edge device's hardware. Hardware health information may include, but is not limited to, the edge device's chassis temperature, CPU temperature, memory temperature, fan speed, fan status (such as normal, abnormal, or failed), power supply status, input / output power, battery status, CPU power consumption, total power consumption, hardware event logs, etc.

[0025] For example, hardware event logs can include logs of power failures, over-temperature warnings, hardware failures, etc.

[0026] For example, out-of-band management interfaces may include, but are not limited to, IPMI (Intelligent Platform Management Interface), Redfish interface, vendor out-of-band interfaces, etc.

[0027] For example, the out-of-band information acquisition module can be integrated with transceiver functions. The data center can directly receive the hardware health information sent by the out-of-band information acquisition module, or the hardware health information can be sent by the out-of-band information acquisition module through an independent transceiver module. That is, the data center receives the hardware health information sent by the out-of-band information acquisition module through the transceiver module. This application does not limit this.

[0028] Step 102: Perform out-of-band monitoring of edge devices based on hardware health information.

[0029] In edge computing scenarios, out-of-band monitoring is a monitoring method that is independent of the main network communication path. It directly collects the operating status, performance indicators, and fault information of edge devices through a dedicated management channel (such as an independent network link, physical serial port, or dedicated management interface) to achieve real-time perception of device status and early warning of anomalies.

[0030] In this application, the data center can use hardware health information to perceive the health status of edge devices in real time and issue anomaly alerts.

[0031] For example, data centers can visualize the hardware health information of edge devices in real time.

[0032] In this embodiment, the data center only needs to establish a communication connection with the out-of-band information acquisition module deployed in the edge cluster to centrally obtain the hardware health information of all edge devices in the cluster. Compared with the method of allocating a public IP address to each edge device and having the data center directly call its out-of-band management interface, this solution significantly reduces the occupation of public IP resources and the number of network connections. Moreover, the out-of-band management interface of the edge device does not need to be directly exposed to the public network, effectively reducing data transmission overhead, communication latency and security exposure. Thus, while ensuring monitoring capabilities, it improves the scalability, security and operation and maintenance efficiency of the system.

[0033] Figure 2 This is a flowchart illustrating an edge device monitoring method provided in another embodiment of this application.

[0034] like Figure 2 As shown, the edge device monitoring method includes: Step 201: Receive hardware health information sent by the out-of-band information acquisition module deployed on the edge cluster.

[0035] In this application, step 201 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0036] Step 202: Based on the hardware health information, obtain the indicator data of each out-of-band monitoring indicator of the edge device.

[0037] In this application, hardware health information may include indicator data of various out-of-band monitoring indicators of the edge device.

[0038] For example, out-of-band monitoring metrics may include, but are not limited to, temperature control metrics such as chassis temperature, CPU temperature, and memory temperature; fan-related metrics such as fan speed and fan status; power supply-related metrics such as power supply status, input / output power, and battery status; and power management metrics such as CPU power consumption and overall system power consumption.

[0039] For example, indicator data may include, but is not limited to, indicator values, collection time, device identification, etc.

[0040] Step 203: Generate and display the device monitoring page for the edge device based on the indicator data of each out-of-band monitoring indicator.

[0041] In this application, the device monitoring page may include indicator data of various out-of-band monitoring metrics, such as the current edge device's chassis temperature, CPU temperature, memory temperature, fan speed, fan status (e.g., normal, abnormal, malfunctioning, etc.), power supply status, input and output power, battery status, CPU power consumption, and total power consumption.

[0042] In this application, the data center can provide a web console, which can be used to display out-of-band monitoring metrics for all edge devices, such as displaying a device monitoring page for a single device.

[0043] In some embodiments, the device monitoring page can be generated based on the indicator data of each out-of-band monitoring indicator: the out-of-band monitoring indicator curve can be generated based on the indicator values ​​of the out-of-band monitoring indicator at each time point within the most recent preset time period, and the device monitoring page can be generated based on the indicator values ​​and curves in the indicator data of the out-of-band monitoring indicator.

[0044] The recent preset duration is associated with out-of-band monitoring metrics. This means that for different out-of-band monitoring metrics, the metric values ​​at various points in time within the recent corresponding preset duration can be obtained. For example, for chassis temperature, the chassis temperature at each hour within the most recent week can be obtained; for CPU temperature, the CPU temperature at each hour within the most recent two days can be obtained.

[0045] As can be seen, the device monitoring page can include not only the data of each out-of-band monitoring indicator, but also the curves of each out-of-band monitoring indicator.

[0046] Therefore, by generating corresponding curves based on the out-of-band monitoring indicators at various time points within a preset time period associated with them, and combining these curves with real-time indicator data to construct a device monitoring page, users can intuitively observe the historical trends of various out-of-band monitoring indicators. This not only enhances the visualization capabilities of the hardware status evolution process, but also supports maintenance personnel in timely identifying potential abnormal patterns and providing early warnings of possible failures, thereby significantly improving the proactive maintenance capabilities of edge devices and the reliability of the system.

[0047] In this embodiment, the data center obtains out-of-band monitoring data of various edge devices based on hardware health information. Based on this data, it generates and displays a device monitoring page for the edge devices. This allows users to intuitively and in real-time view various out-of-band monitoring indicators and their health status through the device monitoring page, significantly improving the visualization of device status and operational transparency. This not only reduces the complexity of fault location and diagnosis but also enhances the maintainability and management efficiency of edge devices, providing strong support for the efficient operation and maintenance of large-scale edge infrastructure.

[0048] Figure 3 This is a flowchart illustrating an edge device monitoring method provided in another embodiment of this application.

[0049] like Figure 3 As shown, the edge device monitoring method includes: Step 301: Receive hardware health information sent by the out-of-band information acquisition module deployed on the edge cluster.

[0050] In this application, step 301 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0051] Step 302: Determine the hardware health of the edge devices under the target monitoring level based on the hardware health information of the edge devices under the target monitoring level of the data center.

[0052] To improve the flexibility of edge device monitoring, this application allows the data center to monitor hardware health at different monitoring levels.

[0053] In this application, the target monitoring level may include, but is not limited to, the regional level, the edge cluster level, etc.

[0054] For example, a data center can monitor hardware health at the region level; that is, a data center can monitor hardware health at the region level. For example, at the region level, a region can include multiple edge clusters, and the region here can refer to a geographical area.

[0055] For example, a data center can monitor the hardware health at the edge cluster level, meaning that a data center can monitor the hardware health at different edge cluster levels according to the edge cluster dimension.

[0056] In this application, for a target monitoring level, the hardware health of each edge device under that monitoring level can be determined based on the hardware health information of each edge device under that monitoring level.

[0057] For example, the hardware health score of an edge device can be used to characterize the health of the edge device; the higher the hardware health score, the healthier the hardware of the edge device.

[0058] Step 303: Determine the hardware health of the target monitoring level based on the hardware health of the edge devices under the target monitoring level.

[0059] In this application, the average hardware health of each edge device under the target monitoring level can be used as the hardware health of the target monitoring level.

[0060] For example, the hardware health of the target monitoring level can be used to characterize the overall hardware health of edge devices under the target monitoring level.

[0061] For example, at the edge cluster level, the average hardware health of all edge devices in an edge cluster can be used as the hardware health of that edge cluster.

[0062] For example, at the regional level, the hardware health of each edge cluster in a region can be determined first, and then the average hardware health of each edge cluster can be used as the hardware health of the region.

[0063] Step 304: Generate and display the grouped view of the target monitoring level based on the hardware health of the target monitoring level.

[0064] In this application, the grouped view of the target monitoring level may include the hardware health of the target monitoring level, the number of edge devices under the target health level, etc.

[0065] For example, a data center can display the hardware health status of a target monitoring level through a web console.

[0066] Optionally, the data center can display hardware health at different monitoring levels through a web console. For example, users can switch between displaying hardware health at different monitoring levels by triggering a toggle control, or display hardware health at different monitoring levels simultaneously on the same page.

[0067] In this embodiment, the hardware health of edge devices at the target monitoring level is determined based on their hardware health information. The overall hardware health of edge devices at the target monitoring level is further determined, and corresponding grouped views are generated for visualization. This allows users to intuitively and in real-time grasp the overall health status of any monitoring level, from devices to regions or edge clusters. This not only enables a structured presentation of multi-level hardware health status but also significantly improves the global perception, fault prediction efficiency, and collaborative management effectiveness of maintenance personnel for large-scale edge infrastructure.

[0068] In one embodiment of this application, the following method can also be used to perform out-of-band monitoring of edge devices based on hardware health information: the number of edge devices with abnormal events can be determined based on the hardware health information of the edge devices monitored by the data center, an overview view can be generated based on the number of devices, and the overview view can be displayed.

[0069] For example, a data center can identify anomalies in the hardware health information of each monitored edge device to determine whether there are any abnormal events on the edge device, and thus count the number of edge devices with abnormal events.

[0070] For example, the percentage of edge devices with abnormal events can be determined based on the ratio between the number of edge devices with abnormal events and the number of edge devices being monitored.

[0071] For example, the overview view may include the number of edge devices with abnormal events, the percentage of edge devices with abnormal events in the total number of monitored edge devices, etc.

[0072] In this embodiment, by analyzing the hardware health information of the monitored edge devices, the number of devices with abnormal events is accurately counted. Based on this statistical result, an intuitive overview view is generated and displayed, allowing users to clearly understand the number of problematic devices among all edge devices. This not only simplifies the anomaly monitoring and management process but also supports the operations and maintenance team in quickly identifying key areas requiring attention and taking timely and effective countermeasures. This improves overall operational efficiency and response speed, ensuring system stability and reliability. Furthermore, it enhances the predictability of potential risks, helping to prevent wider service interruptions or data loss, further ensuring business continuity and data security.

[0073] In some embodiments of this application, the data center may display one or more of the above-mentioned device monitoring page, group view, and overview view, and there is no limitation thereto.

[0074] In this embodiment, the data center can provide a unified, cross-regional hardware-level visual monitoring interface.

[0075] To improve the speed of anomaly response, in one embodiment of this application, the hardware health information may include the indicator data of each out-of-band monitoring indicator, or it may be implemented in the following way: based on the hardware health information, out-of-band monitoring of edge devices is performed: according to the processing strategy corresponding to the sensor type associated with the out-of-band monitoring indicator, the indicator data is processed for anomaly identification, the anomaly information of the out-of-band monitoring indicator is obtained, and based on the anomaly information, a standardized event corresponding to the out-of-band monitoring indicator is generated, an alarm judgment is performed based on the standardized event to generate alarm information, and the alarm information is sent to the target object.

[0076] For example, abnormal information may include, but is not limited to, abnormal indicator values, abnormal hardware events, etc.

[0077] For example, sensor types may include continuous sensors, discrete sensors, etc., and the sensor type associated with out-of-band monitoring indicators may refer to the sensor type used for acquiring the indicator data of the out-of-band monitoring indicators. For example, the out-of-band information acquisition module can acquire the indicator data of out-of-band monitoring indicators collected by the sensors.

[0078] For example, temperature, voltage, and fan speed are collected using continuous sensors, while power supply and hard drive faults are collected using offline sensors.

[0079] For example, for indicator data collected by continuous sensors, sampling jitter can be smoothed (e.g., by moving average), and threshold tables obtained from SDR (Sensor Data Record) can be used to determine if limits are exceeded. For instance, a standardized event can be generated and an alarm output only when the indicator value in the indicator data exceeds the warning or critical threshold.

[0080] For example, for indicator data collected by offline sensors, since the collected indicator data includes event codes, the event codes can be parsed to obtain hardware events. The hardware events can be prioritized according to the severity rules of the hardware events, such as Critical, Warning, Info, etc. Based on the priority of the labels, abnormal hardware events can be identified and converted into standardized events.

[0081] Here, Info indicates that the hardware event is an informational event, meaning that the hardware event is normal. Among the three priorities of Critical, Warning, and Info, Critical indicates the highest severity of the hardware event and has the highest priority, followed by Warning and then Info.

[0082] Since the hardware event log can record events of various hardware components of the edge device (such as chassis, CPU, fan, etc.), the log content can optionally be structured and parsed, such as timestamps, sensor types, event directions, etc., to obtain the events recorded in the log, classify them according to the event anomaly level, filter out abnormal events, and convert abnormal events into standardized events.

[0083] For example, alarm judgment based on standardized events may include at least one of the following: temperature over-limit alarm, fan stall alarm, power failure alarm, and log event alarm (such as power failure, over-temperature, hardware protection trigger, etc.).

[0084] For example, alarm information can be pushed to the target object through one or more channels. For example, the target object can be maintenance personnel or other personnel, etc., without limitation.

[0085] In this embodiment, by processing strategies corresponding to the sensor types associated with out-of-band monitoring indicators, anomaly identification processing is performed on the indicator data to obtain anomaly information of the out-of-band monitoring indicators. Based on the anomaly information, standardized events corresponding to the out-of-band monitoring indicators are generated. This transforms noisy and poorly structured hardware health information into a standardized, stable event stream that can be directly used for edge monitoring, effectively eliminating a large amount of invalid or redundant data. This improves the data effectiveness of the monitoring system, allowing the data center to focus more on critical hardware anomalies. Furthermore, alarm judgment based on standardized events can provide rapid hardware-level diagnostic evidence at the first moment when edge device failures occur. This not only significantly reduces the difficulty of locating and troubleshooting edge device downtime and other failures, but also enhances the system's proactive early warning capabilities and operational response efficiency, providing strong support for the high availability and reliability of large-scale edge infrastructure.

[0086] To reduce transmission overhead, in one embodiment of this application, the data center can provide acquisition tasks to the out-of-band information acquisition module of the edge cluster, enabling the out-of-band information acquisition module to obtain hardware health information based on the acquisition tasks.

[0087] For example, the data center can obtain monitoring requirement information of edge devices, generate target acquisition tasks for the out-of-band information acquisition module based on the monitoring requirement information, and send the target acquisition tasks to the out-of-band information acquisition module so that the out-of-band information acquisition module can collect information from the target hardware according to the acquisition strategy.

[0088] The target acquisition task can be used to instruct the acquisition strategy for the target hardware in each piece of hardware of the edge device. For example, the target acquisition task may include the hardware identifier of the target hardware, the acquisition strategy, etc., and the acquisition strategy may include, but is not limited to, the acquisition content, the acquisition time, etc.

[0089] For example, the monitoring requirement information of edge devices can be obtained based on the detected monitoring trigger operations of maintenance personnel on the target hardware.

[0090] For example, a data center can provide a custom monitoring page that lists hardware and monitorable metrics for each piece of hardware for operations and maintenance personnel to select. If an operations and maintenance personnel trigger a monitoring operation on the target hardware on this page, the monitoring requirement information for the target hardware of the edge device can be obtained. This monitoring requirement information may include the hardware to be monitored, the corresponding metrics to be monitored, etc., and a target data collection task can be generated based on this monitoring requirement information.

[0091] For example, if network instability or poor network quality is detected between the current data center and the out-of-band information acquisition module, the monitoring requirements of the edge devices can be determined, including reducing data transmission volume, i.e., collecting information from some hardware. When the monitoring requirements include reducing data transmission volume, the target hardware can be identified from among the various hardware components. Based on the performance parameters of the target hardware, a collection strategy for the target hardware can be determined. Based on the target hardware and the collection strategy, a target collection task is generated.

[0092] For example, identifying the target hardware from among the various hardware components may include selecting hardware with a historical failure rate exceeding a preset threshold as the target hardware, or selecting hardware with higher priority as the target hardware.

[0093] For example, if the target hardware is a fan, the data collection strategy could be to collect the fan's average speed, maximum speed, minimum speed, etc. during peak hours.

[0094] For example, the target acquisition task can be actively sent by the data center to the out-of-band information acquisition module, or it can be sent by the data center after receiving a pull request from the out-of-band information acquisition module; this application does not limit this. Therefore, the out-of-band information acquisition module can pull the acquisition task and report the acquisition results only when needed, reducing invalid data transmission.

[0095] In this embodiment, a data collection task for the target hardware is dynamically generated based on the monitoring needs of the edge devices, and the task is sent to the out-of-band information collection module. This module then performs data collection according to a customized collection strategy, enabling it to accurately adapt to the hardware configuration, business scenarios, or operation and maintenance strategies of different edge devices. This avoids the waste of resources caused by indiscriminate collection, significantly improving the personalization, flexibility, and targeting of monitoring. Furthermore, it optimizes the utilization efficiency of network bandwidth, storage, and computing resources, laying the foundation for building an efficient and intelligent edge operation and maintenance system.

[0096] To implement the above embodiments, this application also proposes an edge device monitoring method. Figure 4 This is a flowchart illustrating an edge device monitoring method according to another embodiment of this application. The edge device monitoring method is executed by an out-of-band information acquisition module deployed in an edge cluster.

[0097] like Figure 4 As shown, the edge device monitoring method includes: Step 401: The out-of-band information acquisition module deployed in the edge cluster collects information from each hardware component of the edge device in the edge cluster through the out-of-band management interface to obtain hardware health information.

[0098] In this application, the out-of-band information acquisition module, out-of-band management interface, hardware health information, etc. are explained and described in the above embodiments, and will not be repeated here.

[0099] In this application, for any edge cluster with an out-of-band information acquisition module deployed, the edge devices in the edge cluster are located on the same local area network. The out-of-band information acquisition module can call the out-of-band management interface within the LAN (Local Area Network) to obtain the hardware health information of all edge devices in the edge cluster.

[0100] For example, the out-of-band information acquisition module can obtain hardware health information from the BMC (Baseboard Management Controller) of the edge device in the edge cluster by calling the out-of-band management interface. The BMC is an independent microcontroller integrated on the edge device, such as a server motherboard, and is responsible for hardware status monitoring, power management, event logging, etc.

[0101] For example, the BMC includes a sensor subsystem, a logging module, and a power management module. The BMC collects data from temperature sensors, voltage detection chips, and fan speed detection sensors in the sensor subsystem to monitor hardware status. The BMC records event logs through its built-in logging module. The BMC obtains power supply status, input and output power, and battery status through the power management module.

[0102] For example, the out-of-band information acquisition module can periodically collect information from each hardware component of the edge device in the edge cluster through the out-of-band management interface to obtain hardware health information, or it can collect information from each hardware component of the edge device according to the acquisition tasks provided by the data center, etc., without limitation.

[0103] For example, the out-of-band information acquisition module may have a retry mechanism for acquisition exceptions, such as automatically retrying N times when acquisition fails, and generating an acquisition exception event when acquisition is not possible. Here, N is a positive integer, and the value of N can be set according to actual needs, without limitation.

[0104] Therefore, by collecting hardware health information of each hardware device in the edge cluster through the out-of-band information acquisition module, the monitoring range of edge devices is expanded, covering out-of-band information of the hardware layer, and making up for the deficiency that monitoring can only rely on software indicators.

[0105] Step 402: Send hardware health information to the data center so that the data center can perform out-of-band monitoring of the edge devices based on the hardware health information.

[0106] In this application, hardware health information can be encapsulated according to a unified target format to obtain encapsulated data, which can then be sent to the data center via HTTP (HyperText Transfer Protocol) or MQTT (Message Queuing Telemetry Transport).

[0107] Therefore, by encapsulating hardware health information in a unified format and sending it to the data center, the problem of inconsistent information formats output by different equipment manufacturers and different protocols, which prevents the data center from uniformly parsing, displaying and storing the information, can be solved, thus achieving consistent monitoring capabilities for equipment from different manufacturers.

[0108] For example, the out-of-band information acquisition module can be integrated with transceiver functions. The out-of-band information acquisition module can directly send hardware health information to the data center, or the out-of-band information acquisition module can also send hardware health information to the data center through a separate transceiver module. This application does not limit this.

[0109] In this embodiment, the out-of-band information acquisition module deployed on the edge cluster can uniformly aggregate the hardware health information of all edge devices in the cluster. The data can be centrally reported to the data center using only a single public IP address. Compared with the method of allocating a separate public IP address to each edge device and having the data center directly call its out-of-band management interface, this solution significantly reduces the occupation of public IP resources and the number of network connections. Moreover, the out-of-band management interface of the edge device does not need to be directly exposed to the public network, effectively reducing data transmission overhead, communication latency, and security exposure. Thus, while ensuring monitoring capabilities, it improves the scalability, security, and operation and maintenance efficiency of the edge cluster.

[0110] In addition, it can ensure that hardware health information can still be reliably reported in unstable edge network environments, solving the problem that the data center cannot access edge devices through out-of-band management interfaces during failures.

[0111] Figure 5 This is a flowchart illustrating an edge device monitoring method provided in another embodiment of this application.

[0112] like Figure 5 As shown, the edge device monitoring method includes: Step 501: The out-of-band information acquisition module deployed in the edge cluster collects information from each hardware component of the edge device in the edge cluster through the out-of-band management interface to obtain hardware health information.

[0113] In this application, step 501 can be implemented in any of the embodiments of this application, so it will not be described in detail here.

[0114] Step 502: Based on the hardware health information, determine the target abnormal event of the edge device.

[0115] In this application, the target abnormal event can be associated with the target monitoring metric of the edge device, or in other words, the target abnormal event refers to an abnormal event of the target monitoring metric of the edge device. For example, if the target abnormal event is that the chassis temperature exceeds the alarm threshold, then the target monitoring metric is the chassis temperature.

[0116] In some embodiments, in order to reduce network bandwidth usage, not all out-of-band monitoring metrics data are reported to the data center, but abnormal events of out-of-band monitoring metrics are reported to the data center.

[0117] For example, hardware health information may include indicator data for various out-of-band monitoring metrics. For any out-of-band monitoring metric, anomaly detection can be performed on the metric value in the indicator data. If an anomaly is detected in the value of a target monitoring metric among the out-of-band monitoring metrics, a target anomaly event for the target monitoring metric can be generated. Therefore, performing anomaly detection on the indicator data and generating a target anomaly event when an anomaly is detected can improve the accuracy of anomaly event detection.

[0118] For example, hardware health information may include hardware event logs. Since the hardware event logs record events of various hardware components of the edge device (such as the chassis, CPU, and fans), the hardware event logs in the hardware health information can be parsed to obtain the log parsing results. The target abnormal event can then be extracted from the log parsing results based on the event's abnormality level. Therefore, by parsing the hardware event logs and extracting the target abnormal event, the accuracy of abnormal event identification can be improved.

[0119] Step 503: In response to the target abnormal event meeting the abnormal reporting conditions of the target monitoring indicators, send the target abnormal event to the data center.

[0120] To improve the accuracy of anomaly reporting and avoid frequent reporting that increases transmission costs, this application allows for the pre-setting of anomaly reporting conditions for each out-of-band monitoring indicator. If the target anomaly event meets the anomaly reporting conditions of the target monitoring indicator, the target anomaly event is then sent to the data center.

[0121] For example, if the CPU temperature exceeds 85°C, a temperature overrun event is generated. If no similar event has been reported in the last 5 minutes, the CPU temperature overrun event can be sent to the data center.

[0122] It should be noted that in this application, the anomaly reporting conditions corresponding to different out-of-band monitoring indicators can be the same or different, and there is no limitation on this.

[0123] In this embodiment, target abnormal events of edge devices are determined based on hardware health information. When a target abnormal event meets the abnormal reporting conditions of the target monitoring indicator, the target abnormal event is sent to the data center. Therefore, reporting is only required when an abnormal event meets preset reporting conditions. This avoids frequently pushing transient jitter, self-recovery events, or low-priority logs to the data center, reducing invalid alarms and data redundancy, minimizing unnecessary event transmission, and significantly reducing uplink bandwidth usage.

[0124] In one embodiment of this application, the following method can also be used to collect information on each piece of hardware in the edge cluster through the out-of-band management interface to obtain hardware health information: the out-of-band information collection module receives the target collection task sent by the data center and, according to the collection strategy, collects information on the target hardware through the out-of-band management interface to obtain hardware health information associated with the target hardware.

[0125] Specifically, a target acquisition task can be used to instruct acquisition strategies for target hardware in various devices. For example, a target acquisition task can be generated by the data center based on monitoring requirements from edge devices.

[0126] For example, the out-of-band information acquisition module can send an acquisition task retrieval request to the data center. Based on the retrieval request, the data center sends the target acquisition task to the out-of-band information acquisition module, thereby actively pulling the target acquisition task from the data center. Alternatively, the target acquisition task can also be actively sent by the data center, without limitation.

[0127] It should be noted that within the same edge cluster, the target acquisition tasks of different edge devices can be the same or different, and there is no limitation on this.

[0128] In this embodiment, the out-of-band information acquisition module performs data acquisition based on the acquisition task issued by the data center and the acquisition strategy indicated by the acquisition task. This enables the out-of-band information acquisition module to accurately adapt to the hardware configuration, business scenarios, or operation and maintenance strategies of different edge devices, thereby avoiding the waste of resources caused by indiscriminate acquisition. It can not only significantly improve the personalization, flexibility, and targeting of monitoring, but also optimize the utilization efficiency of network bandwidth, storage, and computing resources, laying the foundation for building an efficient and intelligent edge operation and maintenance system.

[0129] In one embodiment of this application, hardware health information can be sent to the data center using a retry mechanism.

[0130] For example, hardware health information can be sent to the data center in the following way: if the hardware health information fails to be sent, the next resend time of the hardware health information can be determined according to the backoff and retry policy. If the resend time is reached, the hardware health information is resent to the data center until the hardware health information is successfully sent or the maximum number of times can be sent is reached.

[0131] For example, backoff and retry strategies may include exponential backoff or incremental backoff.

[0132] Taking exponential backoff as an example, after the first transmission of hardware health information fails, retrying can be delayed by exponentially increasing time intervals (such as 1s, 2s, 4s, 8s, etc.). For example, after the first transmission fails, it can wait 1 second and then retransmit the second time; if the second transmission fails, it can wait 2 seconds and then retransmit the third time; if the third transmission fails, it can wait 4 seconds and then retransmit the fourth time, and so on, until the transmission is successful or the maximum number of retries is reached.

[0133] Taking incremental backoff as an example, after the first transmission of hardware health information fails, retrying can be delayed by incremental time intervals (such as 1s, 3s, 5s, 7s, etc.). For example, after the first transmission fails, it can wait 1 second and then retransmit the second time; if the second transmission fails, it can wait 3 seconds and then retransmit the third time; if the third transmission fails, it can wait 5 seconds and then retransmit the fourth time, and so on, until the transmission is successful or the maximum number of retries is reached.

[0134] For example, if hardware health information fails to be sent, the hardware health information can be stored in a data buffer queue for retransmission, thereby preventing data loss.

[0135] In this embodiment of the application, if the transmission of hardware health information fails, the reliability and accuracy of the transmission of hardware health information can be improved by retransmitting the hardware health information using a backoff and retry strategy.

[0136] To implement the above embodiments, this application also proposes an edge device monitoring system. Figure 6 This is a schematic diagram of the structure of an edge device monitoring system provided in an embodiment of this application.

[0137] like Figure 6 As shown, the edge device monitoring system 600 includes: a data center 610 and an edge cluster 620 with an out-of-band information acquisition module 621 deployed; The data center 610 is used to execute the method of any of the above-described data center side embodiments, and the out-of-band information acquisition module 621 is used to execute the method of any of the above-described out-of-band information acquisition module side embodiments.

[0138] For example, the out-of-band information acquisition module 621 can acquire information from each hardware component of the edge device 622 in the edge cluster 620 through the out-of-band management interface to obtain hardware health information.

[0139] In this embodiment, the out-of-band information acquisition module deployed in the edge cluster can uniformly aggregate the hardware health information of all edge devices in the cluster. The data center can centrally obtain the hardware health information of all edge devices in the cluster by establishing a communication connection with the out-of-band information acquisition module deployed in the edge cluster through a single public IP address. Compared with the method of allocating a public IP address to each edge device and having the data center directly call its out-of-band management interface, this method can significantly reduce the occupation of public IP resources and the number of network connections. Moreover, the out-of-band management interface of the edge device does not need to be directly exposed to the public network, which effectively reduces data transmission overhead, communication latency and security exposure surface. Thus, while ensuring monitoring capabilities, it improves the scalability, security and operation and maintenance efficiency of the edge cluster.

[0140] To facilitate understanding of the edge device monitoring scheme in this application embodiment, the following explanation uses the example of an out-of-band information acquisition module pushing hardware health information to a data center via an independent data push module. Here, the data push module can be equivalent to the independent transceiver module mentioned in the above embodiments.

[0141] Figure 7 This is a schematic diagram illustrating the interaction between an out-of-band information acquisition module of multiple edge clusters and a data center, as provided in an embodiment of this application.

[0142] like Figure 7 As shown, the data center performs out-of-band monitoring of edge devices in M ​​edge clusters: edge cluster A, edge cluster B, edge cluster C, ..., edge cluster M. Each edge cluster contains multiple edge devices, and each edge device is equipped with an out-of-band information acquisition module and a corresponding data push module.

[0143] For any edge cluster, the out-of-band information acquisition module can collect hardware health information of each edge device in the edge cluster and push it to the data center through the data push module. The data center receives and stores the hardware health information, then filters the hardware health information, visualizes the hardware health information, and makes real-time alarm judgments based on the filtered data.

[0144] For example, the specific process of the data center filtering hardware health information can be found in the above embodiments. The data center performs anomaly identification processing on the indicator data according to the processing strategy corresponding to the sensor type associated with the out-of-band monitoring indicator, and generates standardized events based on the identified anomaly information. Therefore, it will not be described in detail here.

[0145] For example, the filtered data can be equivalent to the standardized events in the above embodiments.

[0146] To achieve the above embodiments, this application also proposes an edge device monitoring device. Figure 8 This is a schematic diagram of the structure of an edge device monitoring device provided in an embodiment of this application.

[0147] like Figure 8 As shown, the edge device monitoring device 800 includes: The receiving module 810 is used to receive hardware health information sent by the out-of-band information acquisition module deployed in the edge cluster; wherein, the hardware health information is obtained by the out-of-band information acquisition module through the out-of-band management interface by collecting information from each hardware of the edge devices in the edge cluster. The monitoring module 820 is used to perform out-of-band monitoring of edge devices based on hardware health information.

[0148] Optionally, the monitoring module 820 is used for: Based on hardware health information, obtain the indicator data of each out-of-band monitoring indicator of the edge device; Based on the data of each out-of-band monitoring indicator, generate and display the device monitoring page for the edge device.

[0149] Optionally, the monitoring module 820 is used for: Based on the index values ​​of out-of-band monitoring indicators at various time points within the most recent preset time period, a curve of the out-of-band monitoring indicators is generated; where the preset time period is related to the out-of-band monitoring indicators. Generate a device monitoring page based on the indicator values ​​and graphs in the indicator data.

[0150] Optionally, the monitoring module 820 is used for: Determine the hardware health status of edge devices under the target monitoring level based on their hardware health information. Determine the hardware health of the target monitoring level based on the hardware health of the edge devices under the target monitoring level; Based on the hardware health status of the target monitoring level, generate and display the grouped view of the target monitoring level.

[0151] Optionally, the monitoring module 820 is used for: Based on the hardware health information of the edge devices monitored by the data center, determine the number of edge devices with abnormal events among the monitored edge devices; Generate and display an overview view based on the number of devices.

[0152] Optionally, the hardware health information includes indicator data for each out-of-band monitoring metric. The monitoring module 820 is used for: Based on the processing strategy corresponding to the sensor type associated with the out-of-band monitoring indicators, anomaly identification processing is performed on the indicator data to obtain the abnormal information of the out-of-band monitoring indicators. Based on the anomaly information, generate standardized events corresponding to out-of-band monitoring indicators; Alarms are determined based on standardized events to generate alarm information; Send an alarm message to the target object.

[0153] Optionally, the device may further include: The acquisition module is used to acquire monitoring requirement information from edge devices; The generation module is used to generate target acquisition tasks for the out-of-band information acquisition module based on monitoring requirements; the target acquisition task is used to indicate the acquisition strategy for the target hardware in each hardware. The sending module is used to send the target acquisition task to the out-of-band information acquisition module, so that the out-of-band information acquisition module can acquire information from the target hardware according to the acquisition strategy.

[0154] It should be noted that the explanation of the aforementioned edge device monitoring method embodiment on the data center side also applies to the edge device monitoring device of this embodiment, so it will not be repeated here.

[0155] In this embodiment, the data center only needs to establish a communication connection with the out-of-band information acquisition module deployed in the edge cluster to centrally obtain the hardware health information of all edge devices in the cluster. Compared with the method of allocating a public IP address to each edge device and having the data center directly call its out-of-band management interface, this solution significantly reduces the occupation of public IP resources and the number of network connections. Moreover, the out-of-band management interface of the edge device does not need to be directly exposed to the public network, effectively reducing data transmission overhead, communication latency and security exposure. Thus, while ensuring monitoring capabilities, it improves the scalability, security and operation and maintenance efficiency of the system.

[0156] To achieve the above embodiments, this application also proposes an edge device monitoring device. Figure 9 This is a schematic diagram of the structure of an edge device monitoring device provided in another embodiment of this application.

[0157] like Figure 9 As shown, the edge device monitoring device 900 includes: The out-of-band information acquisition module 910 is used to collect information from each hardware device in the edge cluster through the out-of-band management interface and obtain hardware health information. The sending module 920 is used to send hardware health information to the data center so that the data center can perform out-of-band monitoring of edge devices based on the hardware health information.

[0158] Optionally, the transmitting module 920 is used for: Based on hardware health information, target abnormal events of edge devices are identified; among them, target abnormal events are correlated with target monitoring indicators of edge devices. In response to a target anomaly event meeting the anomaly reporting conditions of the target monitoring indicators, the target anomaly event is sent to the data center.

[0159] Optionally, the hardware health information includes indicator data for each out-of-band monitoring metric, and the sending module 920 is used for: Perform anomaly detection on indicator values ​​in the indicator data; In response to the detection of abnormal indicator values ​​of target monitoring indicators among out-of-band monitoring indicators, a target abnormality event of the target monitoring indicator is generated.

[0160] Optionally, the transmitting module 920 is used for: Parse the hardware event log in the hardware health information and obtain the log parsing results; Extract the target exception event from the log parsing results.

[0161] Optionally, the out-of-band information acquisition module 910 is used for: Receive the target acquisition task sent by the out-of-band information acquisition module from the data center; wherein, the target acquisition task is used to indicate the acquisition strategy for the target hardware in each device; According to the acquisition strategy, information is collected from the target hardware through the out-of-band management interface to obtain hardware health information associated with the target hardware.

[0162] Optionally, the transmitting module 920 is used for: In response to the failure to send hardware health information, the next retransmission time of the hardware health information is determined according to the backoff retry policy; In response to the arrival of the retransmission time, retransmit the hardware health information to the data center until the hardware health information is successfully transmitted or the maximum number of transmissions is reached.

[0163] It should be noted that the explanation of the aforementioned edge device monitoring method embodiment on the out-of-band information acquisition module side also applies to the edge device monitoring device of this embodiment, so it will not be repeated here.

[0164] In this embodiment, the out-of-band information acquisition module deployed on the edge cluster can uniformly aggregate the hardware health information of all edge devices in the cluster. The data can be centrally reported to the data center using only a single public IP address. Compared with the method of allocating a separate public IP address to each edge device and having the data center directly call its out-of-band management interface, this solution significantly reduces the occupation of public IP resources and the number of network connections. Moreover, the out-of-band management interface of the edge device does not need to be directly exposed to the public network, effectively reducing data transmission overhead, communication latency, and security exposure. Thus, while ensuring monitoring capabilities, it improves the scalability, security, and operation and maintenance efficiency of the edge cluster.

[0165] According to embodiments of this application, this application also provides an electronic device, a readable storage medium, and a computer program product.

[0166] Figure 10 A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of this application is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0167] like Figure 10 As shown, device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 1002 or loaded from storage unit 1008 into RAM (Random Access Memory) 1003. RAM 1003 may also store various programs and data required for the operation of device 1000. The computing unit 1001, ROM 1002, and RAM 1003 are interconnected via bus 1004. I / O (Input / Output) interface 1005 is also connected to bus 1004.

[0168] Multiple components in device 1000 are connected to I / O interface 1005, including: input unit 1006, such as keyboard, mouse, etc.; output unit 1007, such as various types of monitors, speakers, etc.; storage unit 1008, such as disk, optical disk, etc.; and communication unit 1009, such as network card, modem, wireless transceiver, etc. Communication unit 1009 allows device 1000 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0169] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the various methods and processes described above, such as edge device monitoring methods. For example, in some embodiments, the edge device monitoring method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 1008. In some embodiments, part or all of the computer program may be loaded and / or installed on device 1000 via ROM 1002 and / or communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the edge device monitoring method described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform an edge device monitoring method by any other suitable means (e.g., by means of firmware).

[0170] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0171] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0172] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0173] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0174] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0175] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is established by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service ecosystem, addressing the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0176] According to an embodiment of this application, this application also provides a computer program product that, when an instruction processor in the computer program product is executed, performs the edge device monitoring method proposed in the above embodiments of this application.

[0177] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.

[0178] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A method for monitoring edge devices, comprising: Receive hardware health information sent by an out-of-band information acquisition module deployed in the edge cluster; wherein, the hardware health information is obtained by the out-of-band information acquisition module through the out-of-band management interface by collecting information from each hardware of the edge devices in the edge cluster; Based on the hardware health information, out-of-band monitoring is performed on the edge device.

2. The method as described in claim 1, wherein, The step of performing out-of-band monitoring of the edge device based on the hardware health information includes: Based on the hardware health information, obtain the indicator data of each out-of-band monitoring indicator of the edge device; Based on the indicator data of each out-of-band monitoring indicator, generate and display the device monitoring page of the edge device.

3. The method as described in claim 2, wherein, The step of generating a device monitoring page for the edge device based on the indicator data of each out-of-band monitoring indicator includes: Based on the index values ​​of the out-of-band monitoring index at each time point within the most recent preset time period, a curve of the out-of-band monitoring index is generated; wherein, the preset time period is associated with the out-of-band monitoring index; The device monitoring page is generated based on the indicator values ​​and the curve in the indicator data.

4. The method of claim 1, wherein, The step of performing out-of-band monitoring of the edge device based on the hardware health information includes: Based on the hardware health information of the edge devices under the target monitoring level of the data center, determine the hardware health of the edge devices under the target monitoring level; The hardware health of the target monitoring level is determined based on the hardware health of the edge devices under the target monitoring level. Based on the hardware health of the target monitoring level, generate a grouped view of the target monitoring level and display the grouped view.

5. The method of claim 1, wherein, The step of performing out-of-band monitoring of the edge device based on the hardware health information includes: Based on the hardware health information of the edge devices monitored by the data center, determine the number of edge devices among the monitored edge devices that have abnormal events; An overview view is generated based on the number of devices, and the overview view is displayed.

6. The method of claim 1, wherein, The hardware health information includes indicator data for various out-of-band monitoring metrics. The step of performing out-of-band monitoring of the edge device based on the hardware health information includes: Based on the processing strategy corresponding to the sensor type associated with the out-of-band monitoring indicator, anomaly identification processing is performed on the indicator data to obtain the abnormal information of the out-of-band monitoring indicator. Based on the anomaly information, a standardized event corresponding to the out-of-band monitoring indicator is generated; Alarm judgment is performed based on the standardized events to generate alarm information; The alarm information is sent to the target object.

7. The method according to any one of claims 1-6, further comprising: Obtain the monitoring requirements information of the edge devices; Based on the monitoring requirement information, a target acquisition task is generated for the out-of-band information acquisition module; wherein, the target acquisition task is used to indicate the acquisition strategy for the target hardware among the hardware. The target acquisition task is sent to the out-of-band information acquisition module so that the out-of-band information acquisition module can acquire information from the target hardware according to the acquisition strategy.

8. A method for monitoring edge devices, comprising: The out-of-band information acquisition module deployed in the edge cluster collects information from each hardware component of the edge devices in the edge cluster through the out-of-band management interface to obtain hardware health information; The hardware health information is sent to the data center so that the data center can perform out-of-band monitoring of the edge device based on the hardware health information.

9. The method of claim 8, wherein, Sending the hardware health information to the data center includes: Based on the hardware health information, a target abnormal event of the edge device is determined; wherein, the target abnormal event is associated with a target monitoring indicator of the edge device; In response to the target anomaly event meeting the anomaly reporting conditions of the target monitoring indicator, the target anomaly event is sent to the data center.

10. The method of claim 9, wherein, The hardware health information includes indicator data for each out-of-band monitoring metric. Determining the target abnormal event of the edge device based on the hardware health information includes: Anomaly detection is performed on the indicator values ​​in the aforementioned indicator data; In response to detecting an abnormal value of the target monitoring indicator among the out-of-band monitoring indicators, a target abnormality event of the target monitoring indicator is generated.

11. The method of claim 9, wherein, The step of determining the target abnormal event of the edge device based on the hardware health information includes: The hardware event log in the hardware health information is parsed to obtain the log parsing results; Extract the target exception event from the log parsing results.

12. The method of claim 8, wherein, The step of collecting information on the hardware of each edge device in the edge cluster through the out-of-band management interface to obtain hardware health information includes: The system receives a target acquisition task from the out-of-band information acquisition module sent by the data center; wherein the target acquisition task is used to indicate an acquisition strategy for the target hardware in each device. According to the acquisition strategy, information is collected from the target hardware through the out-of-band management interface to obtain the hardware health information associated with the target hardware.

13. The method of claim 8, wherein, Sending the hardware health information to the data center includes: In response to the failure to send the hardware health information, the next retransmission time of the hardware health information is determined according to the backoff and retry strategy; In response to the arrival of the retransmission time, the hardware health information is retransmitted to the data center until the hardware health information is successfully transmitted or the maximum number of transmissions is reached.

14. An edge device monitoring system, comprising: Data centers and edge clusters deployed with out-of-band information acquisition modules; The data center is used to perform the method according to any one of claims 1-7, and the out-of-band information acquisition module is used to perform the method according to any one of claims 8-13.

15. An edge device monitoring device, comprising: The receiving module is used to receive hardware health information sent by the out-of-band information acquisition module deployed in the edge cluster; wherein, the hardware health information is obtained by the out-of-band information acquisition module through the out-of-band management interface by collecting information from each hardware of the edge devices in the edge cluster. The monitoring module is used to perform out-of-band monitoring of the edge device based on the hardware health information.

16. An edge device monitoring device, comprising: The out-of-band information acquisition module is used to acquire information from each hardware device in the edge cluster through the out-of-band management interface and obtain hardware health information. The sending module is used to send the hardware health information to the data center so that the data center can perform out-of-band monitoring of the edge device based on the hardware health information.

17. An electronic device comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-13.

18. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-13.

19. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1-13.