Server failure monitoring system
By deploying data acquisition terminals and edge servers in the server room for initial data analysis, only abnormal data is transmitted to the central server for fault identification, and the data is visualized on the monitoring terminal. This solves the problem of low efficiency in fault analysis in the server room, and enables rapid fault location and reduced response latency.
Patent Information
- Application Number
- CN202511255616.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-09-04
AI Technical Summary
In server rooms, existing technologies cannot effectively and automatically link data from different monitoring systems, resulting in low efficiency in fault analysis and long response delays.
Data is collected from the server using a data acquisition terminal, and initial analysis is performed through an edge server. Only abnormal data is sent to the central server for fault identification, and then visualized on the monitoring terminal. This progressive data processing workflow reduces hardware deployment costs and improves the accuracy of fault identification.
It improves fault monitoring efficiency, reduces fault response delay, reduces hardware deployment costs, and allows maintenance personnel to directly locate faulty devices, avoiding switching back and forth between multiple systems.
Smart Images

Figure CN120803854B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of server technology, and in particular to a server fault monitoring system. Background Technology
[0002] With the rapid development of artificial intelligence, cloud technology, and big data, server rooms are becoming increasingly larger. In the process of monitoring server room faults, data from different monitoring systems cannot be automatically correlated. It is necessary to switch between multiple monitoring systems and analyze large amounts of data, resulting in low fault analysis efficiency and long response delays. Summary of the Invention
[0003] This invention provides a server fault monitoring system that effectively improves server fault monitoring efficiency and reduces fault response delay.
[0004] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0005] This invention provides a server fault monitoring system, including data acquisition terminals deployed in a server room, edge servers connected to the corresponding area data acquisition terminals, a central server connected to the edge servers, and a monitoring terminal connected to the central server.
[0006] Each data acquisition terminal generates server monitoring data based on the device identifier, environmental monitoring data, operational status monitoring data, log data, and data acquisition time of the monitored server, and sends it to the connected edge server. Each edge server, when an anomaly log exists in the target server's monitoring data, extracts the data to be identified related to the anomaly log and generates an anomaly reporting log for the target monitored server. The central server, upon determining a fault in the target monitored server based on the anomaly reporting log and the data to be identified, generates fault data information and sends it to the monitoring terminal. The monitoring terminal, based on the target device identifier of the target monitored server, determines its server graphical object in the 3D virtual model of the computer room, updates the display status of the server graphical object on the visualization page to a fault state, and simultaneously generates a fault data interaction tag for the server graphical object based on the fault data information.
[0007] The advantages of the technical solution provided by this invention are as follows: It utilizes a data acquisition terminal to collect data from the monitored server, uploads the collected data to an edge server for initial analysis to determine if any anomalies exist, and only sends abnormal data to the central server for fault identification. This prevents useless data transmission from consuming excessive bandwidth, reducing latency caused by data transmission without consuming network resources. After identifying a fault, the central server displays it on the visualization page of the monitoring terminal. By integrating hardware and software information collection and processing through a progressive data processing flow, it not only eliminates the need for multiple deployments of devices with different products and functions, effectively reducing hardware deployment costs, but also improves the accuracy of fault identification. After identifying a fault, its location is displayed on the monitoring terminal, allowing maintenance personnel to directly locate the faulty device without switching between multiple systems, effectively improving fault monitoring and repair efficiency and minimizing fault response latency. Attached Figure Description
[0008] To more clearly illustrate the technical solutions of the present invention or related technologies, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is an exemplary structural diagram of the server fault monitoring system provided by the present invention;
[0010] Figure 2 A schematic diagram of the hardware composition structure applicable to the server fault monitoring system provided by the present invention;
[0011] Figure 3 This is a framework intent for the server fault monitoring system provided by the present invention in an exemplary application scenario. Detailed Implementation
[0012] To enable those skilled in the art to better understand the technical solutions of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. In this specification and the aforementioned drawings, the terms "first," "second," "third," "fourth," etc., are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. The term "exemplary" means "serving as an example, embodiment, or illustration." Any embodiment described herein as "exemplary" is not necessarily to be construed as superior to or better than other embodiments.
[0013] Server fault monitoring involves monitoring device operating status, collecting environmental parameters, and issuing fault warnings. These processes are implemented through independently operating systems, such as basic sensor monitoring systems, log analysis systems, and network monitoring platforms. Each system collects data from different dimensions, resulting in inconsistent data formats and a lack of automatic data correlation between different monitoring systems, leading to low fault analysis efficiency. Furthermore, alarm information is scattered across different platforms. The process is cumbersome and complex, requiring time to analyze logs and switching between multiple systems to compare data. This results in long waiting times from fault occurrence to repair and significant response delays. For example, when a server fails, maintenance personnel must first locate the device in the asset management system, then analyze the error code in the log system, and finally confirm the connection status through network monitoring—a highly inefficient process.
[0014] In view of this, to solve the problems existing in related technologies, the present invention utilizes a data acquisition terminal to collect environmental and software data of the monitored server, uploads the collected data to an edge server for initial analysis to determine whether anomalies exist, and only sends abnormal data to a central server for fault identification. After fault identification, it is displayed on the visualization page of the monitoring terminal. Through a progressive data processing flow, the efficiency of server fault monitoring is effectively improved, and fault response delay is minimized. After introducing the technical solution of the present invention, various non-limiting embodiments of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments.
[0015] First, please see Figure 1 , Figure 1 This is an exemplary structural diagram of the server fault monitoring system provided in this embodiment. This embodiment may include the following:
[0016] A server fault monitoring system may include at least a data acquisition terminal 1, an edge server 2, a central server 3, and a monitoring terminal 4. Among these, for example... Figure 2As shown, the number of data acquisition terminals 1 is the same as the number of monitored servers. They are deployed on each monitored server in the server room requiring fault monitoring. Due to resource constraints of edge servers 2, multiple edge servers 2 can be deployed according to regions. Data acquisition terminals 1 located in the region of an edge server 2 interact with the edge server 2 in that region. For example, each rack can be configured with one edge server 2, and the data acquisition terminals 1 on the edge server 2 in that rack can be connected to the edge server 2 in that rack via wired or wireless connections. Each edge server 2 is connected to the central server 3 via wired or wireless connections. The monitoring terminal 4 is a server configured with a display, and it is connected to the central server 3 via wired or wireless connections. In other words, each data acquisition terminal 1 is connected to the edge server 2 in its corresponding region, each edge server 2 is connected to the central server 3, and the central server 3 is connected to the monitoring terminal 4.
[0017] Each data acquisition terminal 1 collects hardware and software monitoring data from the server it resides on. The server where data acquisition terminal 1 resides is defined as the monitored server. The software monitoring data includes the monitored server's operational status monitoring data and log data. The operational status monitoring data includes, but is not limited to, remotely collected system version, CPU (Central Processing Unit) information and utilization rate, memory information and utilization rate, storage information and utilization rate, important processes, user information, network configuration (including IP (Internet Protocol Address), subnet mask, gateway, etc.) and connection status. The log data includes, but is not limited to, system logs, security logs, important application logs, BMC (Baseboard Management Controller) logs, and operation logs. Of course, the software monitoring data may also include other data that users need to pay attention to, and this invention does not impose any limitations on this. The collection frequency of the software monitoring data can be customized according to actual needs, such as collecting key data (such as CPU, memory, various logs, etc.) every 2 seconds and collecting secondary data every 10 seconds. For ease of management, time and device identification information can be added after each software monitoring data collection. The time information can be, for example, a timestamp of the data collection date, and the device identification information is the unique identifier of the monitored server. Hardware monitoring data includes at least environmental monitoring data, which may include at least temperature, humidity, and vibration data, and can be collected through sensors. Similarly, time and device identification information can be added after each software monitoring data collection. Therefore, each data acquisition terminal 1 generates server monitoring data based on the device identification of the corresponding monitored server, environmental monitoring data, operational status monitoring data, log data, and data collection time, and sends the server monitoring data to the connected edge server 2.
[0018] Specifically, for each edge server 2, whenever an edge server 2 receives server monitoring data sent by a data acquisition terminal 1, it will identify whether there is abnormal data in the server monitoring data. Abnormal data can be identified by identifying whether there are abnormal alarm logs or error logs. For example, it can be identified by identifying whether abnormal keywords such as IERR (Internal ERRor), CATERR (Catastrophic Error), CE (Correctable Error), and UCE (Uncorrectable Error) appear in the logs, whether there are warnings for temperature exceeding preset temperature thresholds, whether humidity is outside the appropriate range, and abnormal vibration warnings. When an abnormality is identified in the received server monitoring data, it is defined as target server monitoring data for ease of description, and the monitored server corresponding to the target server monitoring data is defined as the target monitored server. In other words, when abnormal logs exist in the target server's monitoring data, the data to be identified related to these abnormal logs is extracted from the target server's monitoring data. This data includes the necessary information to determine whether a fault has occurred during the monitoring of the server and its internal devices. An abnormal reporting log for the target monitored server is then generated. This log can be generated according to a pre-set log format, and it must include at least the target monitored server's device identifier (defined as the target device identifier), an anomaly type field (e.g., whether the software monitoring data is abnormal or the hardware monitoring data is abnormal), the event that caused the anomaly, and key parameters indicating the anomaly. Finally, the abnormal reporting log and the data to be identified are packaged and sent to the central server 3. Through initial data analysis by the edge server 2, a large amount of normal logs and irrelevant data can be filtered out, preventing useless data transmission from consuming excessive bandwidth.
[0019] In this system, the central server 3 can be, for example, a server in the main control room of a server room, and the monitoring terminal 4 can be a display terminal in the main control room. Whenever the central server 3 receives an anomaly report log and data to be identified from an edge server 2, it determines whether the target monitored server has a fault based on the anomaly report log and the data to be identified. For example, it can pre-train a model with fault identification capabilities using a language model or any neural network model, using log information and fault handling methods from recent years as training sample data, combined with deep learning algorithms. Then, the anomaly report log and the data to be identified are input into this model, and the model's output determines whether a fault has occurred. More simply, any large language model from a related technology can be fine-tuned directly using instructions to enable fault identification. When a fault is determined to exist in the target monitored server based on the anomaly report log and the data to be identified, fault data information is generated. This fault data information includes at least the target device identification information and fault type of the target monitored server, and then the fault data information is sent to the monitoring terminal 4.
[0020] The monitoring terminal 4 constructs a corresponding virtual room (defined as a 3D virtual room model) according to the actual layout of the server room. During the construction of the 3D virtual room model, a globally unique device identifier is assigned to each physical device to achieve precise mapping and data association. The device identifier can be, for example, an IP address, a MAC (Media Access Control Address) address, an asset tag, or a unique ID generated by the asset management system (CMDB). The choice of identifier depends on the existing operation and maintenance system specifications and the availability of the data source. For example, if the monitoring system manages devices based on IP addresses, then using the IP address as the unique identifier is the most direct choice. During the 3D model creation or import phase, each device model needs to be bound to its corresponding unique identifier. This process is usually completed during the model data preparation phase. It can be done by embedding the identifier in the model file's metadata or by dynamically assigning the device identifier to each graphical object via a program script after the 3D engine loads the model. Therefore, one-click location can be achieved by inputting the target device identifier of the monitored server, that is, determining the server graphical object of the monitored server in the 3D virtual room model. To intuitively locate faulty devices, the display status of server graphical objects on the visualization page can be updated to reflect their fault status. Device status can be distinguished by color: green for normal status, yellow for warning status, and red for fault status. In addition to color differentiation, richer visual styles can be defined for device status, such as different light intensities, transparency, or metallic hues, to further differentiate the status visually. To facilitate maintenance personnel in viewing the cause of the fault, fault data interaction labels are generated for the server graphical objects based on the fault data information. These labels can be displayed to the user by clicking on the server graphical objects. The labels include the cause of the fault and can also include maintenance suggestions, centrally displaying the processed data. The interface has a high degree of visualization, allowing users to easily obtain the necessary fault information from multiple dimensions, and can be used without extensive training.
[0021] In the technical solution provided in this embodiment, a data acquisition terminal collects data from the monitored server, and the collected data is uploaded to an edge server for initial analysis to determine if there are any anomalies. Only data with anomalies is sent to the central server for fault identification, preventing useless data transmission from consuming a large amount of bandwidth. This reduces latency caused by data transmission without consuming network resources. After identifying a fault, the central server displays it on the visualization page of the monitoring terminal. By integrating hardware and software information collection and processing through a progressive data processing flow, not only is it possible to reduce hardware deployment costs by eliminating the need to deploy different products and devices with different functions multiple times, but it also improves the accuracy of fault identification. After identifying a fault, its location is displayed on the monitoring terminal, allowing maintenance personnel to directly locate the faulty device without switching between multiple systems, effectively improving the efficiency of fault monitoring and fault repair, and minimizing fault response latency. The entire process uses an automated processing method, which reduces business incidents caused by missed fault detection and ensures the stability and reliability of the server room.
[0022] The above embodiments do not limit the data acquisition terminal 1 in any way. The present invention also provides an exemplary implementation method, which may include the following:
[0023] The data acquisition terminal 1 includes a hardware sensor group, a computer-readable storage medium, a data processor, a scanner, and a data interface. The data acquisition terminal 1 connects to the corresponding edge server 2 via the data interface. The hardware sensor group includes temperature, humidity, and vibration sensors deployed at the air outlets of each monitored server in the server room. The sensor data is sent to the data processor as environmental perception data. The scanner reads the device tag information of the monitored server and sends the device tag information to the corresponding edge server 2 via the data interface. The computer-readable storage medium stores a computer program. This computer program is injected into the target location of the computer program code on the monitored server, acquiring at least the server's operating status data and various log data. After adding a timestamp and the monitored server's device identifier, it is sent to the corresponding edge server 2. The data processor extracts temperature, vibration, and humidity values from the environmental perception data to obtain the monitored server's environmental monitoring data. After adding a timestamp and the monitored server's device identifier, it is sent to the corresponding edge server 2 via the data interface.
[0024] In this embodiment, as Figure 3As shown, the hardware sensor group of data acquisition terminal 1 may include a temperature probe (for monitoring temperature and whether high temperatures are present), a humidity sensor (for monitoring humidity and whether it is within a suitable range), and a triaxial vibration sensor (for monitoring machine vibration and whether abnormal vibrations are present). The data interface can be a fiber optic network port, compatible with USB 3.0 data transmission. The scanner can scan QR codes or barcodes to automatically read device tag information (including IP, serial number, and maintenance records). The hardware sensor group can collect basic data (temperature, humidity, current) at a custom frequency, such as every 10 seconds. When a parameter exceeds a threshold, the sampling frequency is automatically increased to twice per second. All data is timestamped and device numbered before being uploaded to edge server 2 via dual channels (wired + backup wireless). A computer program on a computer-readable storage medium can be automatically injected into key locations (such as function entry / exit points) of the monitored server software to achieve uninterrupted data acquisition. The data processor performs a filtering process on the raw data collected by the sensors, extracting specific data (temperature, humidity, vibration), and removing unnecessary data before uploading it to edge server 2, reducing response latency through efficient data transmission.
[0025] The above embodiments do not limit the selection method for generating anomaly reporting logs and data to be identified. Based on the above embodiments, the present invention also provides an exemplary implementation method, which may include the following:
[0026] For cases where the abnormal log is a hardware alarm abnormal log, the target device identifier, abnormal alarm log, abnormal detection data, and abnormal alarm time of the monitored server are extracted from the target server monitoring data. An environment log is generated based on the target device identifier, abnormal hardware identifier, abnormal alarm type corresponding to the abnormal alarm log, abnormal detection data, and hardware abnormal alarm time. The unidentified operating status data, target operating status parameters exceeding the target operating status parameter threshold, and their duration are extracted from the target server monitoring data. A device operation log is generated based on the target operating status parameters and their duration, the target device identifier, the hardware abnormal alarm time, and the software identifier. An abnormal reporting log is generated based on the environment log and the device operation log, and the unidentified operating status data and abnormal alarm log are used as the data to be identified.
[0027] For cases where the abnormal log is a software alarm abnormal log, the following are extracted from the target server monitoring data: target device identifier, software abnormal alarm time, software alarm abnormal log, unidentified operating status data at the time of the software abnormal alarm, target operating status parameters exceeding the target operating status parameter threshold and their duration. Based on the target operating status parameters and their duration, target device identifier, software abnormal alarm time, software identifier, and the software alarm type corresponding to the software alarm abnormal log, a device operation log is generated. The target environment monitoring data corresponding to the software abnormal alarm time is extracted from the target server monitoring data, and an environment log is generated based on the target device identifier, hardware identifier, and software abnormal alarm time. An abnormal reporting log is generated based on the environment log and the device operation log, and the unidentified operating status data, software alarm abnormal log, and target environment monitoring data are used as the data to be identified.
[0028] In this embodiment, after merging a large number of normal logs, only one error message is displayed. For abnormal logs, relevant information is extracted, such as timestamps, device numbers, module identifiers, key parameters, important values, and event supplements (optional). This relevant information is used to generate a short log to replace the original log data. For example, the format of the environmental log could be: [2025-08-13T19:51:03.452] [A25-C35] [Hardware HW] [Temperature Alarm TEMP_ALERT] Air outlet 89.2 degrees (threshold 85 degrees). The format of the device operation log could be: [2025-08-13T19:51:03.755][010-A01] [Software SW] [CPU Load] 98% 15s duration. An anomaly reporting log format could be, for example, [2025-08-13T19:51:04.128] [C6669] [HW / SW] [Critical Alarm] Memory temperature 92 degrees + process crash (threshold 80 degrees + PID: 4412). If relevant data needs to be collected based on the fault type, a statistical summary log can also be generated. The format of the statistical summary log could be, for example,: [2025-08-13T19:51:10.000] [Cabinet 10, No. 1] [Statistics COUNT] Anomaly count within 10 minutes: Temperature TEMP (representing temperature sensor) × 3 | Memory MEM (representing memory) × 1 | Network Card NET (representing network card) × 0.
[0029] The above embodiments do not limit how to construct the 3D virtual model of the computer room. The present invention also provides an exemplary construction method for the 3D virtual model of the computer room. Considering that the central server 3 usually adopts a high-performance server with high computing power, this embodiment constructs the 3D virtual model of the computer room through the central server 3, and then sends the constructed 3D virtual model of the computer room to the monitoring terminal 4, which may include the following:
[0030] Central server 3 acquires physical space data and equipment asset information of the server room, determines the hierarchical structure of the model used to organize various entity objects in the server room, and generates unique identification information for each interactive entity object in the server room; based on the hierarchical structure, it performs 3D modeling of the server room according to the physical space data and equipment asset information, and renders the 3D model to obtain the corresponding 3D model of the server room; the target device graphic objects in the 3D model of the server room include at least user data attributes, color configuration items, and interaction labels; it acquires rack management information, chassis management information, motherboard management information, motherboard system information, and hard disk information of the server room, and generates a device status dataset; based on the unique identification information of each interactive entity object, it obtains the device status data of each interactive entity object from the device status dataset, locates the corresponding device graphic object of each interactive entity object in the 3D model of the server room, fills the corresponding device status data into the user data attributes, and updates the corresponding color configuration items to obtain the 3D virtual model of the server room, and sends the 3D virtual model of the server room to the monitoring terminal 4.
[0031] In this embodiment, the physical space data of the 3D virtual model of the computer room includes, but is not limited to, the building structure, dimensions, internal layout, and location and specifications of all fixed facilities of the computer room. This data can be obtained through the building design drawings of the computer room, such as floor plans, elevations, and sections in CAD (Computer-Aided Design) format. In practice, maintenance personnel or modelers need to carefully analyze these CAD drawings to extract key spatial information. By importing DWG or DXF files into 3D modeling software or specialized conversion tools, the basic wall outlines of the computer room can be automatically generated, greatly simplifying repetitive work in the initial stages of modeling. In addition to the building structure, floor information, such as the material, dimensions, and installation method of the anti-static flooring, also needs to be collected, which is crucial for simulating the real environment of the computer room. In some cases, if the CAD drawings are missing or do not match the actual situation, on-site surveys may be necessary, using tools such as laser rangefinders to conduct on-site measurements to ensure data accuracy. After acquiring the physical space data of the data center, detailed information on all IT equipment and related assets is needed, including servers, switches, routers, storage devices, UPS (Uninterruptible Power Supply), precision air conditioners, power distribution cabinets, surveillance cameras, and all other critical equipment deployed within the data center. For each type of equipment, multi-dimensional information needs to be collected: physical attributes such as the equipment's brand, model, precise length, width, and height dimensions, and its specific installation location within the data center (e.g., which rack and which unit it's in); and logical attributes such as the equipment's asset number, IP address, MAC address, and purpose (e.g., web server, database server). This information can typically be obtained from an IT asset management system (CMDB) or network management system. Associating this physical and logical information is fundamental to subsequent equipment status visualization and intelligent positioning functions. For example, during modeling, a unique identifier needs to be assigned to each equipment model. This identifier can correspond to the equipment number or IP address in the asset management system, thereby enabling precise control and data binding of specific equipment in the 3D scene. To ensure the consistency, maintainability, and high performance of the 3D data center virtual model, it is necessary to define several aspects, including model structure, naming conventions, material mapping, and detail levels. First, regarding the model structure, a clear hierarchical structure should be used to organize objects in the scene. For example, the entire data center can be treated as a top-level group, containing subgroups such as building structures, server racks, and environmental equipment. Each server rack group then contains equipment models such as servers and switches. This hierarchical organization not only facilitates management but also benefits subsequent interactive operations and performance optimization. Second, regarding naming conventions, a unified and meaningful naming rule must be established for each interactive object in the scene (such as server racks and servers).For example, names can be formatted as cabinet-001, server-192.168.1.10 to ensure uniqueness and readability. Standardized naming is crucial for enabling rapid device retrieval and data binding. Regarding materials and textures, specifications should clearly define texture sizes, formats (e.g., JPG, PNG), and UV mapping standards to ensure consistency and efficiency when rendering models on the web. Finally, at the level of detail, different precision versions of the model should be created based on the importance of the device and its distance from the camera to balance visual appeal and rendering performance.
[0032] Modeling the computer room's architectural structure is fundamental to constructing a virtual scene. This process is typically completed using professional 3D modeling software. First, the CAD floor plan of the computer room can be imported as a reference base map. Based on this base map, techniques such as polygon modeling or spline extrusion are used to precisely create the basic structure of the computer room, including walls, floors, ceilings, doors, windows, and columns. For example, a 2D outline of the wall can be drawn first, and then height can be assigned using the extrusion modifier to form a 3D wall. For doors and windows, openings can be cut into the walls using Boolean operations, and then pre-made door and window models can be installed. To improve efficiency and accuracy, the array tool can be used to quickly generate repeating structural elements, such as the mesh for anti-static flooring. During the modeling process, it is essential to strictly adhere to the pre-defined dimensional specifications to ensure that the spatial proportions of the virtual computer room are completely consistent with the physical computer room. After completing the basic structural modeling, appropriate materials can be assigned to the model, such as latex paint on the walls and a metallic texture on the floor, to enhance the realism of the scene.
[0033] Since 3D models need to be rendered in real-time in web browsers, their complexity and data volume directly affect system performance and user experience. Therefore, model optimization and lightweighting are necessary to minimize the number of polygons (faces) and texture sizes while maintaining acceptable visual effects. The construction process for the 3D model in the server room may include the following:
[0034] Based on the fact that the spatial proportions of the virtual machine room are the same as those of the physical machine room, a virtual model of the server room's architectural structure is generated according to the physical space data. According to the scale of the 3D virtual machine room model and the fault monitoring requirements, the first type of target equipment is identified through geometric modeling, the second type of target equipment whose geometric details are not displayed, and the third type of target equipment whose geometric model is replaced by textures. According to the type of each equipment in the equipment asset information, the corresponding 3D model of each equipment is performed, and the equipment is placed in the corresponding position of the architectural structure virtual model based on the model hierarchy, thus obtaining the 3D model of the server room.
[0035] First, the principle of "good enough" should be followed during the modeling stage to avoid creating unnecessary geometric details. For example, low-poly models can be used for devices far from the camera; for small or minor devices, their structure can be simplified appropriately. Second, textures can be used to represent model details. For example, ventilation holes, buttons, and labels on a cabinet can be achieved with a high-precision transparent texture map instead of modeling complex geometry, which greatly reduces the number of polygons. In addition, normal maps can be used to simulate the bumps and depressions of surfaces, improving the texture of the model without increasing geometric complexity. Before exporting the model, optimization tools in the modeling software can be used to automatically reduce the number of polygons in the model. Finally, regarding textures, the size of the textures should be reasonably controlled, avoiding the use of excessively large texture files, and textures should be compressed to reduce memory usage and loading time.
[0036] After constructing the 3D virtual model of the server room, it is necessary to achieve data synchronization and state mapping between the physical world and the 3D virtual model. This involves associating the server's real-time device status data (such as normal operation, warnings, and faults) with the corresponding device graphical objects in the 3D virtual model, and dynamically rendering this association through intuitive visual encoding (color differentiation). To associate business data (such as device ID, IP address, and status information) with device graphical objects (such as a mesh representing the server), the custom attribute mechanism provided by the 3D engine can be utilized. In this embodiment, each device graphical object has a user data attribute specifically for storing user-defined data. The user data attribute can be a regular JavaScript object, which is not used by the internal logic of the 3D engine. Therefore, it can safely store any application-related data without conflicting with the engine's functionality. This avoids directly modifying the attributes of device graphical object instances, preventing potential naming conflicts and compatibility issues that may arise from future engine upgrades. By populating user data attributes, each device graphical object carries rich business information. When locating a device graphical object, it is only necessary to traverse the device graphical objects in the scene graph and check the corresponding attribute values in its user data attributes. To improve operational efficiency and situational awareness, distinct and easily distinguishable colors are assigned to different device states (normal, warning, fault). Device states are dynamically rendered on a 3D model using an intuitive color-coded method, allowing operations personnel to quickly identify devices requiring attention in complex data center layouts. In this implementation, these colors can be defined as global constants or configuration items within the application for unified management and modification. This definition method makes the code more readable and maintainable. When color adjustments are needed, only this one place needs to be modified for global effect.
[0037] After the central server 3 generates a 3D virtual model of the computer room, it can send it to the monitoring terminal 4. The monitoring terminal 4 needs to achieve real-time visualization of the device status. It can obtain real-time data through polling or server push, traversing the received device status data list. For each device in the list, its unique identifier is used to find the corresponding model object in the 3D virtual scene using methods such as `get Objects By Property` (command). Once the target object is found, the device's status information (e.g., "status": "warning") is parsed, and the visual representation of the model object is updated according to preset rules. For example, the color attribute of the model material can be modified, or a material indicating a warning status can be switched. Simultaneously, the status field in the model's user data attributes can also be updated to maintain data consistency.
[0038] Based on the three-dimensional virtual model of the computer room constructed in the above embodiments, the process of updating the status of server graphical objects and generating fault data interaction tags may include:
[0039] Obtain the fault color value corresponding to the fault status, modify the color value of the server graphic object's color configuration item to the fault color value; extract the faulty device, faulty equipment parameters, and fault time from the fault data information, and fill the faulty device, faulty equipment parameters, and fault time into the interactive label to obtain the fault data interactive label.
[0040] Furthermore, this invention also supports the autonomous selection of displaying all or some of the information of all or some devices: When the monitoring terminal 4 receives a request to display the operating status, it obtains the operating status parameters to be displayed and the display conditions, such as displaying all devices with a temperature higher than 40 degrees Celsius or displaying devices with a storage utilization rate exceeding 60%. If the operating status parameters to be displayed include the identifier of the device to be displayed, it determines the corresponding target graphic object in the three-dimensional computer room virtual model according to the identifier of the device to be displayed, and obtains the target operating status data corresponding to the operating status parameters to be displayed from the user data attributes of the target graphic object. If the target operating status data meets the display conditions, it modifies the color configuration item of the target graphic object according to the display conditions. If the operating status parameters to be displayed do not include the identifier of the device to be displayed, it sequentially traverses the user data attributes of each device graphic object in the three-dimensional computer room virtual model, obtains the operating status data corresponding to the operating status parameters to be displayed, and determines each target graphic object that meets the display conditions, and modifies the color configuration item of each target graphic object according to the display conditions.
[0041] As can be seen from the above, this embodiment uses three-dimensional visualization to easily obtain all the information of the device from multiple dimensions, centrally display the processed data, and has a high degree of interface visualization, which is conducive to improving fault response efficiency.
[0042] Furthermore, to reduce fault response delay, based on the above embodiments, the present invention also provides a method for dynamically adjusting data compression rate and transmission priority according to network conditions, which may include the following:
[0043] Edge server 2 acquires first-type network parameter data of the communication links between itself and each data acquisition terminal 1, and second-type network parameter data of the communication links between itself and the central server 3; determines a first network state based on the first-type network parameter data, and determines a second network state based on the second-type network parameter data; matches a corresponding first data compression rate based on the first network state, and sends the first data compression rate to the data acquisition terminal 1 to compress the server monitoring data according to the first data compression rate; matches a corresponding second data compression rate based on the second network state, and compresses the data to be identified according to the second data compression rate; if the current bandwidth utilization is lower than a preset threshold, it transmits the highest priority data according to the data type transmission priority; if the data to be transmitted includes multiple data types, it splits the data to be transmitted into data segments according to the data type transmission priority, and transmits the highest priority data segment first; if the current network is interrupted, it stores the data to be transmitted in the local target storage space.
[0044] The data compression ratio is inversely proportional to the network condition; that is, the better the network condition, the lower the compression ratio. Network parameters include bandwidth, latency, and packet loss rate. The worse the network condition (low bandwidth, high latency), the higher the compression ratio, and vice versa. Preset bandwidth and latency thresholds (e.g., bandwidth thresholds: 10Mbps, 50Mbps; latency thresholds: 50ms, 100ms) are used to compare real-time data with the current network conditions and determine a matching compression ratio (e.g., 90%, 70%). For example, when bandwidth < 10Mbps or latency > 100ms, a compression ratio of over 90% reduces data size, prioritizing the transmission of critical alarm information; when bandwidth > 50Mbps and latency < 50ms, a compression ratio of around 75% ensures complete data transmission. Depending on the actual usage scenario, lossless compression algorithms can be used, reducing size by at least 75%. Furthermore, data types can be prioritized for transmission based on actual network conditions; for example, alarm data can occupy bandwidth first, while basic data can be transmitted with a delay. Furthermore, it also supports offline caching, meaning that data can be stored locally for 72 hours when the network is interrupted.
[0045] Furthermore, based on the above embodiments, the present invention also provides a multi-dimensional alarm and alarm push implementation method, such as... Figure 3 As shown, it may include the following:
[0046] Central server 3 determines the fault urgency level, fault device type, fault frequency, and fault safety level based on fault data information, and generates multi-dimensional alarm information based on these parameters. It then matches authorized user terminals in the alarm information push database and pushes the multi-dimensional alarm information to those terminals. Finally, it matches alarm display terminals and alarm notification terminals in the alarm information push database, controls the display of the multi-dimensional alarm information on the visualization page of the alarm display terminal, and provides alarm notifications on the alarm notification terminal.
[0047] The system allows for setting alarm levels based on urgency: Emergency, Important, Minor, and Alert. Emergency alarms indicate a risk of core system services or critical data leakage, such as server crashes causing complete business interruption or core database failures rendering it inaccessible; immediate action is required. Important alarms indicate service quality degradation requiring urgent repair, such as network bandwidth overload causing a surge in service latency, CPU (Central Processing Unit) utilization consistently exceeding 90% impacting business performance, or server overheating leading to overall performance degradation. Minor alarms indicate situations not yet affecting business but requiring further action, such as disk space below 10% requiring cleanup or non-critical service malfunctions. Alert alarms indicate potential risks or status indicators, such as CPU utilization approaching 80% requiring monitoring or log file size exceeding expected growth rates. Hardware alarms can be set by device type, including power failures, CPU / memory / hard drive anomalies, and abnormal fan speeds, such as hard drive failures or unstable power supply voltage. System alarms involve operating system or software anomalies, such as system crashes, unresponsive services, and insufficient resources. Network alarms: Cover all network-related issues, such as connection interruptions, high latency, and abnormal packet loss rates. Security alarms: Address potential security threats, such as virus attacks, abnormal logins, and distributed denial-of-service attacks. Storage alarms: Issues with all storage-related components, such as insufficient disk space, redundant RAID card failures, and file system errors. Environmental alarms: Related to the physical environment, such as excessively high server temperatures and abnormal heat dissipation. Allocation by frequency: High-frequency alarms (continuous / frequent occurrences): Triggered repeatedly within a short period, possibly due to system design flaws or long-term resource shortages, such as CPU / memory usage consistently exceeding thresholds, frequent network interface packet loss, or latency fluctuations. Medium-frequency alarms (periodic / regular occurrences): Triggered at fixed intervals (e.g., daily / weekly), possibly related to peak business periods or scheduled tasks, such as database connection pool exhaustion during daily peak periods or scheduled backup task failures for some reason. Low-frequency alarms (occasional / random occurrences): Low probability of occurrence, but may imply potential risks, such as unpatched vulnerabilities discovered during security scans or occasional server restarts. Sudden bursts of alarms within a short period: A large number of alarms appearing in a short time are usually triggered by sudden events, such as a surge in resource usage due to a virus intrusion or a large number of transactions being blocked due to database table locking. Irregular and difficult-to-reproduce intermittent alarms: Alarms appear and disappear intermittently, making troubleshooting difficult, such as randomly appearing error codes in logs. Alarms can be set according to their security level: Top Secret Alarms: Highly sensitive alarms involving core system security or core business data, such as any alarm in the core data center or an intrusion alarm on a highly sensitive data center server. Confidential Alarms: Alarms that affect business continuity or contain sensitive operational data; leakage may pose a risk to business competition, such as alarms related to the storage of customer data on servers.Secret-level alarms: These are routine maintenance alarms. Leakage of these alarms may affect system stability but pose no direct security threat, such as alarms for server resource overload or service response delays.
[0048] After generating corresponding alarm information according to the aforementioned multi-dimensional alarm generation rules, corresponding push notifications can be made based on user permissions. An alarm information push database is pre-configured, including permission management for different users. Each user is assigned a position, rank, and permissions. Based on this database, alarms can be pushed to users with corresponding permissions according to urgency, device type, frequency of occurrence, and security level. For example, problems in the core data center can only be pushed to users with core permissions, and memory problems only need to be pushed to the hardware interface. To further improve the effectiveness of alarms, this embodiment employs multiple information push methods, such as an A / B redundancy approach to ensure timely handling of important alarms. Authorized user terminals can be, for example, mobile apps that push top-secret, confidential, urgent, and important (customizable) alarm notifications in real time, supporting viewing detailed alarm and fault information, viewing 3D positioning views, and accessing camera permissions.
[0049] Furthermore, monitoring terminal 4 can also be configured with a web-based console. This console supports displaying data center information according to permissions, filtering devices based on multiple conditions, real-time updates of all information, and customized information display. Alarm notifications can be delivered via voice and mobile devices. For top-secret, confidential, urgent, and important (customizable) alarms, both SMS notifications and voice calls can be provided simultaneously. Additionally, the visualization interface of monitoring terminal 4 includes a device details panel and an alarm handling dashboard. The device details panel displays customized device locations, IP addresses, and other information. Clicking on any device allows viewing real-time parameter curves (temperature, load, network traffic), maintenance history, and a topology diagram of associated devices. The alarm handling dashboard can sort pending tasks by urgency, prioritizing critical alarms with prominent highlighting. Alarms that have not been processed for a long time will receive repeated notifications. It provides standard handling process guidance, including processing suggestions and historical processing procedures, offering quick troubleshooting methods. It generates alarm work orders in real-time, reminding administrators to handle them promptly. Once completed, the work order can be closed, ensuring rapid response and quick processing.
[0050] As can be seen from the above, this embodiment analyzes alarm push from multiple dimensions through permission management and precise push, and delivers alarms accurately and in multiple ways. This not only ensures the effectiveness of alarms, but also prevents data leakage and effectively improves security.
[0051] To enable those skilled in the art to clearly understand the technical solution of the present invention, the present invention also provides exemplary implementation methods, which may include the following:
[0052] In a hard drive failure scenario, when data acquisition terminal 1 detects abnormal vibration of the monitored server's hard drive exceeding a set threshold via a vibration sensor, the server failure monitoring system automatically correlates the device's logs and finds that the number of remapped sectors exceeds the threshold. The push layer simultaneously sends an alarm to the mobile app and web console, including the device location (based on a 3D virtual model of the data center), suggested checks (hard drive module), a list of associated devices (other nodes in the same storage cluster), and historical procedures and suggestions for handling the failure (such as replacing the hard drive). The push target is the storage interface. Maintenance personnel replace the hard drive according to the prompts, and the system automatically updates the maintenance record and clears the alarm.
[0053] In software failure scenarios, the central server identifies anomaly reports containing the keyword "CE" in the logs. Through the identified data, it determines that the number of memory CE alarms for a specific memory module in the core data center exceeds a preset threshold within a short period, and identifies the server device to which this memory belongs as CA-1-54688. Anomaly information, including alarm details, location information, memory information, historical processing flow, and suggestions, is pushed to the administrator. A pending work order is generated in real time, and the push is sent to the hardware interface of the core data center. The data center administrator immediately brings the same type of memory to the data center for repair. After repair, the alarm is deactivated and the pending work order is closed.
[0054] Based on the above embodiments, the central server 3 will receive multi-source data. In order to improve the data processing efficiency of the central server, this embodiment also provides an implementation method for multi-source data fusion analysis, which may include the following:
[0055] First, the device's physical parameters (temperature, vibration), network status (IP, traffic), and software logs (error codes) are aligned along a timeline. Then, the multi-dimensional time data is calibrated using a dynamic time warping algorithm, and the log data are automatically correlated using a sliding time window.
[0056] The protocol used for global time synchronization is determined based on the scenario's synchronization error requirements. If the error budget is ≤1ms, the central server 3, hardware sensor group, edge server 2, and monitoring terminal 4 will uniformly use the PTP (Precision Time Protocol) protocol across the entire network. If the error budget is ≤10ms, NTP (Network Time Protocol) v4 can be used to maintain time synchronization. For the three types of nodes—sensors, edge server 2, and central server 3—data is uniformly encapsulated in a "four-tuple": Logs: Event ID, level, text, local timestamp; Device information: Model, firmware version, MAC, location coordinates; Sensor data: Sample value, dimension, unit, sampling period; Network information: Egress bandwidth, RTT, packet loss rate, link layer type. Furthermore, a 64-bit "synchronized timestamp" field with a precision of 100ns can be added to the header of the four-tuple. Each sensor can send a data frame to the edge server every 20ms via a UDP (User Datagram Protocol) / DTLS (Datagram Transport Layer Security) tunnel. The edge server can send compressed packets to the central server in batches every 50ms via an HTTP / 2 (Hypertext Transfer Protocol version 2) persistent connection. The edge server can pre-configure a local cache with a memory circular queue. When the queue overflows, local DTW (Dynamic Time Warping) pre-alignment is triggered, pushing only the calibrated timestamp upstream, reducing the load on the central server. The central server uses its local PTP discipline clock as the absolute time axis. For log data, it extracts "event level + keyword hash". For sensor data, it uses "first-order difference of sampled values". For network data, it uses "RTT (Round-Trip Time) jitter". For device information, it uses "firmware version number" to identify sudden changes caused by hot upgrades. The window length is set to 2 seconds, the step size to 500 ms, and z-score normalization is performed on each sequence. Then, the shortest regular path is found using Sakoe-Chiba (band constraint strategy) constraints (100 ms bandwidth), outputting "time shift Δt" and "frequency drift slope k" with a precision of 10 µs. The original timestamps are mapped to the global axis using Δt and k to generate a "uniform timestamp". Correction records are written back to the "time alignment table" (TTL (Time To Live) 24h) for data traceability.
[0057] For 500ms sliding window log association operations: the window length is predefined as 500ms, the step size is 100ms, meaning windows overlap by 400ms, and the left boundary of the window equals the current global time – 500ms. The following association rules are set according to priority from high to low: ① Time: Uniform timestamp falls within the window; ② Event: Log event ID matches the sensor trigger flag; ③ Device: Use the same device ID or MAC address; ④ Network: Same link layer type, RTT difference < 5ms. If multiple logs point to the same event within the same 500ms window, the one with the closest "uniform timestamp" to the center of the window is selected. If a log is missing, but both sensor and network information indicate the event, the log field is completed using interpolation, and "Trust Level = Low" can be marked.
[0058] As can be seen from the above, this embodiment, by highly unifying and integrating all data, facilitates rapid fault identification, shortens fault handling time, and effectively reduces the possibility of false alarms.
[0059] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0060] The server fault monitoring system provided by this invention has been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the embodiments above are only intended to help understand the method and core ideas of this application. It should be noted that those skilled in the art can make various improvements and modifications to this application without departing from its principles, and these improvements and modifications also fall within the protection scope of this invention.
Claims
1. A server fault monitoring system, characterized in that, This includes data acquisition terminals deployed in the server room, edge servers connected to the corresponding area data acquisition terminals, a central server connected to each edge server, and a monitoring terminal connected to the central server. Each data acquisition terminal generates server monitoring data based on the device identifier, environmental monitoring data, operational status monitoring data, log data, and data acquisition time of the monitored server, and sends it to the connected edge server. Each edge server extracts the data to be identified related to the abnormal logs from the target server monitoring data when there are abnormal logs in the target server monitoring data, and generates an abnormal reporting log for the target monitored server. When the central server determines that the target monitored server has a fault based on the anomaly reporting log and the data to be identified, it generates fault data information and sends it to the monitoring terminal. Obtain physical space data and equipment asset information of the server room, determine the model hierarchy structure used to organize various entity objects in the server room, and generate unique identification information for each interactive entity object in the server room; Based on the hierarchical structure of the model, a 3D model of the server room is created using physical space data and equipment asset information. This 3D model is then rendered to obtain the corresponding 3D model of the server room. The target device graphic objects in the 3D model of the server room include at least user data attributes, color configuration items, and interaction labels. The system acquires server room rack management information, chassis management information, motherboard management information, motherboard system information, and hard drive information to generate a device status dataset. Based on the unique identifier of each interactive entity object, the system retrieves the device status data of each interactive entity object from the device status dataset. The corresponding device graphic objects are then located in the 3D model of the server room, and the corresponding device status data is filled into the user data attributes. Simultaneously, the corresponding color configuration items are updated to obtain a 3D virtual model of the server room. This 3D virtual model of the server room is then sent to the monitoring terminal. The monitoring terminal determines the server graphical object in the 3D virtual model of the monitored server based on the target device identifier, updates the display status of the server graphical object on the visualization page to a fault status, and generates a fault data interaction tag for the server graphical object based on the fault data information. When a request to display the operating status is received, the terminal obtains the operating status parameters to be displayed and the display conditions. If the operating status parameters to be displayed include the device identifier, the terminal determines the corresponding target graphical object in the 3D virtual model based on the device identifier, and obtains the target operating status data corresponding to the operating status parameters to be displayed from the user data attributes of the target graphical object. If the target operating status data meets the display conditions, the terminal modifies the color configuration item of the target graphical object according to the display conditions. If the operating status parameters to be displayed do not include the device identifier, the terminal sequentially traverses the user data attributes of each device graphical object in the 3D virtual model of the computer room to obtain the operating status data corresponding to the operating status parameters to be displayed, determines each target graphical object that meets the display conditions, and modifies the color configuration item of each target graphical object according to the display conditions.
2. The server fault monitoring system according to claim 1, characterized in that, The display status of the server graphical object on the visualization page is updated to a fault status, and fault data interactive tags for the server graphical object are generated based on the fault data information, including: Obtain the fault color value corresponding to the fault status, and modify the color value of the color configuration item of the server graphical object to the fault color value; Extract the faulty device, faulty equipment parameters, and fault time from the fault data information, and fill the faulty device, faulty equipment parameters, and fault time into the interactive label to obtain the fault data interactive label.
3. The server fault monitoring system according to claim 1, characterized in that, Based on the aforementioned model hierarchy, a 3D model of the server room is performed according to the physical space data and the equipment asset information, including: Based on the fact that the space ratio of the virtual machine room is the same as that of the physical machine room, a virtual model of the building structure of the server room is generated according to the physical space data. Based on the scale of the three-dimensional computer room virtual model and the fault monitoring requirements, the first type of target equipment is identified through geometric modeling, the second type of target equipment whose geometric details are not displayed, and the third type of target equipment whose geometric model is replaced by textures. Based on the equipment type of each device in the equipment asset information, a corresponding 3D model is created for each device, and based on the model hierarchy, each device is placed in the corresponding position of the building structure virtual model to obtain the 3D model corresponding to the server room.
4. The server fault monitoring system according to any one of claims 1 to 3, characterized in that, The data acquisition terminal includes a hardware sensor group, a computer-readable storage medium, a data processor, a scanner, and a data interface; The data acquisition terminal is connected to the corresponding edge server through the data interface; The hardware sensor group includes temperature sensors, humidity sensors, and vibration sensors deployed at the air outlets of each monitored server in the server room, and sends the sensor data as environmental perception data to the data processor. The scanner reads the device tag information of the monitored server and sends the device tag information to the corresponding edge server through the data interface. The computer-readable storage medium stores a computer program, which is injected into the target location of the computer program code of the monitored server to obtain at least the operating status data and various log data of the monitored server, and sends them to the corresponding edge server after adding a timestamp and the device identifier of the monitored server. The data processor extracts temperature, vibration, and humidity values from the environmental sensing data to obtain environmental monitoring data of the monitored server. After adding a timestamp and the device identifier of the monitored server, it sends the data to the corresponding edge server through the data interface.
5. The server fault monitoring system according to any one of claims 1 to 3, characterized in that, The edge server acquires first-type network parameter data of the communication link between the edge server and each data acquisition terminal, and second-type network parameter data of the communication link between the edge server and the central server. A first network state is determined based on the first type of network parameter data, and a second network state is determined based on the second type of network parameter data. A first data compression ratio is matched according to the first network state, and the first data compression ratio is sent to the data acquisition terminal to compress the server monitoring data according to the first data compression ratio; a second data compression ratio is matched according to the second network state, and the data to be identified is compressed according to the second data compression ratio; wherein, the value of the data compression ratio is inversely proportional to the network state; If the current bandwidth utilization is lower than the preset threshold, the data with the highest priority will be transmitted according to the data type transmission priority. If the data to be transmitted includes multiple data types, the data to be transmitted will be split into multiple data segments according to the data type transmission priority, and the data segment with the highest priority will be transmitted first. If the current network is interrupted, the data to be transmitted will be stored in the local target storage space.
6. The server fault monitoring system according to any one of claims 1 to 3, characterized in that, The anomaly log is a hardware alarm anomaly log, which generates anomaly reporting logs for the target monitored server, including: Extract the target device identifier, hardware alarm and anomaly log, anomaly detection data, and hardware anomaly alarm time from the target server monitoring data, and generate an environment log based on the target device identifier, hardware identifier, the anomaly alarm type corresponding to the hardware alarm and anomaly log, the anomaly detection data, and the hardware anomaly alarm time. Extract the target monitored server's unidentified operating status data, target operating status parameters exceeding the target operating status parameter threshold, and their duration from the target server monitoring data at the time of the hardware anomaly alarm, and generate a device operation log based on the target operating status parameters and their duration, the target device identifier, the hardware anomaly alarm time, and the software identifier; An anomaly reporting log is generated based on the environment log and the device operation log, and the operating status data to be identified and the hardware alarm anomaly log are used as the data to be identified.
7. The server fault monitoring system according to any one of claims 1 to 3, characterized in that, The exception log is a software alarm exception log, which generates exception reporting logs for the target monitored server, including: Extract the target device identifier, software anomaly alarm time, software alarm anomaly log, unidentified operating status data within the software anomaly alarm time, target operating status parameters exceeding the target operating status parameter threshold and their duration from the target server monitoring data, and generate device operation log based on the target operating status parameters and their duration, the target device identifier, the software anomaly alarm time, the software identifier, and the software alarm type corresponding to the software alarm anomaly log; Extract the target environment monitoring data corresponding to the software anomaly alarm time from the target server monitoring data, and generate an environment log based on the target device identifier, hardware identifier, and software anomaly alarm time; An anomaly reporting log is generated based on the environment log and the device operation log, and the operating status data to be identified, the software alarm anomaly log, and the target environment monitoring data are used as the data to be identified.
8. The server fault monitoring system according to any one of claims 1 to 3, characterized in that, The central server determines the fault urgency level, fault device type, fault frequency, and fault safety level based on the fault data information, and generates multi-dimensional alarm information. Based on the multi-dimensional alarm information, an authorized user terminal is matched in the alarm information push database, and the multi-dimensional alarm information is pushed to the authorized user terminal. Based on the multi-dimensional alarm information, the alarm display terminal and alarm notification terminal are matched in the alarm information push database, and the multi-dimensional alarm information is controlled to be displayed on the visualization page of the alarm display terminal and alarm notification is issued on the alarm notification terminal.
Citation Information
Patent Citations
Server monitoring method and device, computer equipment and storage medium
CN109815093A
Server fault detection system and implementation method
CN113704051A
Cited By
Knowledge graph-based server operation fault root cause mining method and system
CN122173327A