Server fault monitoring system

Through the combined architecture of data acquisition terminals, edge servers and central servers, efficient monitoring and rapid response to server failures are achieved, solving the problems of low fault analysis efficiency and response delay in existing technologies and providing an intuitive fault location tool.

CN120803854AActive Publication Date: 2025-10-17LANGCHAO ELECTRONIC INFORMATION IND CO LTD

Patent Information

Application Number
CN202511255616.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-04
Publication Date
2025-10-17
Estimated Expiration
2045-09-04

AI Technical Summary

Technical Problem

In server rooms, existing technologies are unable to effectively and automatically correlate data from different monitoring systems, resulting in low fault analysis efficiency and long response delays.

Method used

A combined architecture of data acquisition terminals, edge servers, and central servers is adopted. Server data is monitored in real time through the data acquisition terminals, the edge servers perform initial analysis and filter abnormal data, the central servers perform fault identification, and the data is visualized on the monitoring terminals to achieve progressive data processing.

Benefits of technology

It improves fault monitoring efficiency, reduces fault response delay, reduces hardware deployment costs, and helps operation and maintenance personnel quickly locate faulty equipment through an intuitive visual interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120803854A_ABST
    Figure CN120803854A_ABST
Patent Text Reader

Abstract

The invention discloses a server fault monitoring system, which relates to the technical field of servers, and is characterized in that a data acquisition terminal deployed in a machine room acquires environment monitoring data and software monitoring data of a server, and the environment monitoring data and the software monitoring data are used as server monitoring data to be sent to a corresponding edge server; and when the edge server detects that the software / hardware monitoring data is abnormal, extracting related abnormal data from the corresponding server monitoring data, and generating an abnormal report log. And the central server determines that the server has a fault according to the exception report log and the related exception data, and generates fault data information. And the monitoring terminal locates a graphic object corresponding to the fault server in the three-dimensional machine room virtual model according to the equipment identifier, updates the display state of the graphic object of the server into a fault state, and generates a fault data interaction label according to the fault data information. According to the invention, the problem of relatively large fault response delay in related technologies can be solved, and the server fault monitoring efficiency can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of servers, and in particular to a server fault monitoring system. BACKGROUND

[0002] With the rapid development of artificial intelligence technology, cloud technology and big data, the scale of server rooms is becoming larger and larger. In the process of monitoring the server room for faults, the data of different monitoring systems cannot be automatically associated, the operation needs to be switched between multiple monitoring systems, and the scale of data to be analyzed is large, which leads to low fault analysis efficiency and long response delay time. SUMMARY

[0003] The present application provides a server fault monitoring system, which effectively improves the server fault monitoring efficiency and reduces the fault response delay.

[0004] To solve the above technical problems, the present application provides the following technical solutions: The present application provides a server fault monitoring system, which comprises a plurality of data acquisition terminals deployed in a server room, a plurality of edge servers connected to the corresponding regional data acquisition terminals, a central server connected to the edge servers, and a monitoring terminal connected to the central server.

[0005] Each data acquisition terminal generates server monitoring data according to the device identifier of the monitored server, the environmental monitoring data, the running state monitoring data, the log data and the data acquisition time, and sends the server monitoring data to the connected edge server. When the target server monitoring data has abnormal logs, the edge server extracts the to-be-identified data related to the abnormal logs from the target server monitoring data and generates abnormal report logs of the target monitored server. When the central server determines that the target monitored server has a fault according to the abnormal report logs and the to-be-identified data, it generates fault data information and sends it to the monitoring terminal. The monitoring terminal determines the server graphic object of the target monitored server in the three-dimensional room virtual model according to the target device identifier of the target monitored server, and updates the display state of the server graphic object in the visualization page to the fault state. At the same time, the monitoring terminal generates a fault data interaction tag of the server graphic object according to the fault data information.

[0006] The technical scheme provided by the application has the advantages that the data acquisition terminal is used to acquire the data of the monitored server, the acquired data is uploaded to the edge server for initial analysis to determine whether there is an exception, only the data with the exception is sent to the central server for fault identification, useless data transmission is prevented from occupying a large amount of bandwidth, on the basis of not occupying network resources, the delay caused by data transmission is reduced, the central server displays the fault on the visual page of the monitoring terminal after identifying the fault, the software and hardware information acquisition and information processing are integrated, through the progressive data processing flow, different product and different function devices do not need to be deployed multiple times, the hardware deployment cost is effectively reduced, the fault identification accuracy is improved, the fault is positioned and displayed on the monitoring terminal after being identified, the operation and maintenance personnel can directly locate the fault device, and it is not necessary to switch among multiple systems, the fault monitoring and fault repair efficiency is effectively improved, and the fault response delay is minimized. BRIEF DESCRIPTION OF DRAWINGS

[0007] In order to more clearly illustrate the technical scheme of the present application or related art, the following will briefly introduce the drawings needed to be used in the embodiments or related art description. Obviously, the drawings in the following description are only some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creating any inventive labor.

[0008] Figure 1 An exemplary structural schematic diagram of the server fault monitoring system provided by the present application; Figure 2 An exemplary structural schematic diagram of the hardware composition structure suitable for the server fault monitoring system provided by the present application; Figure 3 A framework diagram of the server fault monitoring system provided by the present application in an exemplary application scenario. DETAILED DESCRIPTION

[0009] In order to make the person skilled in the art better understand the technical scheme of the present application, the present application will be further described in detail below in combination with the drawings and specific embodiments. In the specification and the above drawings, the terms "first", "second", "third", "fourth" and the like are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. The term "exemplary" means "as an example, embodiment or illustration". Any embodiment illustrated as "exemplary" herein is not necessarily interpreted as superior or better than other embodiments.

[0010] The server fault monitoring process involves device running state monitoring, environment parameter collection, fault early warning, and related technologies are implemented through independent systems such as a basic sensor monitoring system, a log analysis system, and a network monitoring platform. These systems respectively collect data in different dimensions, the data formats between systems are not unified, and the data of different monitoring systems cannot be automatically associated, resulting in low fault analysis efficiency. In addition, alarm information is scattered in different platforms. Since time is needed to analyze logs and switching operations are needed between multiple systems to compare data, the operation is cumbersome and complex, which results in a long waiting time from fault occurrence to fault repair and large response delay. For example, when a server fails, an operation and maintenance personnel needs to first find the device location in an asset management system, then analyze the error code in a log system, and finally confirm the connection state through a network monitoring, the whole process is inefficient.

[0011] In view of this, the present application is used to solve the problems existing in the related art. The environment data and software data of the monitored server are collected by the data collection terminal, the collected data is uploaded to the edge server for initial analysis to determine whether there is an anomaly, only the abnormal data is sent to the central server for fault identification, and after the fault is identified, it is displayed on the visualization page of the monitoring terminal. Through the progressive data processing process, the server fault monitoring efficiency is effectively improved, and the fault response delay is minimized. After introducing the technical solution of the present application, various non-limiting embodiments of the present application will be described in detail in combination with the drawings and specific embodiments.

[0012] First, please see Figure 1 , Figure 1 An exemplary structural schematic diagram of the server fault monitoring system provided by the present embodiment can include the following contents: The server fault monitoring system can at least include a data collection terminal 1, an edge server 2, a central server 3, and a monitoring terminal 4. As shown in Figure 2 The data collection terminal 1 is the same as the number of monitored servers, which is deployed on each monitored server in the server room that needs to be monitored for faults. Limited by the resources of the edge server 2, multiple edge servers 2 can be deployed according to the area, and the data collection terminal 1 located in the area of the edge server 2 interacts with the edge server 2 in the area, such as one edge server 2 is configured for each rack, and the data collection terminal 1 of the edge server 2 on the rack is connected to the edge server 2 of the rack, which can be connected by wire or wireless. Each edge server 2 is connected to the central server 3 by wire or wireless, and the monitoring terminal 4 is a server configured with a display, which is connected to the central server 3 by wire or wireless. That is, each data collection terminal 1 is connected to the edge server 2 in the corresponding area, each edge server 2 is connected to the central server 3, and the central server 3 is connected to the monitoring terminal 4.

[0013] Wherein, each data collection terminal 1 collects the software and hardware monitoring data of the server where it is located, the server where the data collection terminal 1 is located is defined as the monitored server, the software monitoring data includes the running state monitoring data and the log data of the monitored server, the running state monitoring data includes but is not limited to the remote collection system version, CPU (Central Processing Unit) information and occupancy rate, memory information and occupancy rate, storage information and occupancy rate, important processes, user information, network configuration (including IP (Internet Protocol Address) address, subnet mask, gateway, etc.) and connection state, the log data includes but is not limited to system log, security log, important application log, BMC (Baseboard Management Controller) log, operation log. Of course, the software monitoring data can also include other data required by users to pay attention to, and the present application does not make any limitation on this. The collection frequency of the software monitoring data can be customized according to actual needs, such as collecting the data (such as CPU, memory, various logs, etc.) that needs to be paid attention to every 2s, and collecting the secondary data every 10s. In order to facilitate management, time information and device identification information can be added after each collection of software monitoring data, the time information can be, for example, the time stamp added when the data is collected, and the device identification information is the unique identification information of the monitored server. The hardware monitoring data at least includes environmental monitoring data, the environmental monitoring data can at least include temperature data, humidity data, vibration data, which can be collected through sensors, and similarly, time information and device identification information can be added after each collection of software monitoring data. As can be seen, each data collection terminal 1 generates server monitoring data according to the device identification, environmental monitoring data, running state monitoring data, log data and data collection time of the corresponding monitored server, and sends the server monitoring data to the connected edge server 2.

[0014] Wherein, for each edge server 2, the edge server 2 will identify whether there is abnormal data in the server monitoring data sent by a data collection terminal 1 whenever receiving the server monitoring data, and the abnormal data can be determined by identifying whether there is an abnormal alarm log or an error log, such as by identifying whether there is an IERR (Internal ERRor), CATERR (Catastrophic Error), CE (Correctable Error), UCE (Uncorrectable Error) and other abnormal keywords in the log, whether the temperature exceeds the preset temperature threshold, whether the humidity is not within the appropriate range, and whether there is an abnormal vibration warning. When it is identified that the received server monitoring data is abnormal, in order to facilitate description, it is defined as target server monitoring data, and the monitored server corresponding to the target server monitoring data is defined as a target monitored server. That is, when there is an abnormal log in the target server monitoring data, the to-be-identified data related to the abnormal log is extracted from the target server monitoring data, the to-be-identified data includes data necessary for judging whether a fault occurs in the server and its internal equipment, and an abnormal report log of the target monitored server is generated. The abnormal report log can be generated according to a pre-set log format, which at least includes a field of device identification (defined as target device identification) of the target monitored server, an abnormal type field (such as software monitoring data abnormality or hardware monitoring data abnormality), an event of abnormal occurrence and a key parameter reflecting the abnormality. Finally, the abnormal report log and the to-be-identified data are packaged and sent to the central server 3. Through the initial data analysis of the edge server 2, a large amount of normal logs and irrelevant data can be filtered out to prevent useless data transmission from occupying a large amount of bandwidth.

[0015] Wherein, the central server 3 may be a server in the main control room of the server room, and the monitoring terminal 4 may be a display terminal in the main control room. The central server 3 will determine whether the target monitored server has a fault according to the abnormal report log and the to-be-identified data whenever receiving the abnormal report log and the to-be-identified data sent by an edge server 2. For example, a language model or any kind of neural network model can be used as training sample data by using log information and fault handling methods in recent years, combined with a deep learning algorithm to train a model with fault identification capability, and then the abnormal report log and the to-be-identified data are input into the model to determine whether a fault occurs according to the model output result. More simply, any large language model in related technology can be directly fine-tuned to have fault identification capability. When it is determined that the target monitored server has a fault according to the abnormal report log and the to-be-identified data, fault data information is generated, which at least includes target device identification information and fault type of the target monitored server, and then the fault data information is sent to the monitoring terminal 4.

[0016] The monitoring terminal 4 constructs a corresponding virtual machine room (defined as a three-dimensional machine room virtual model) according to the actual layout of the server room. In constructing the three-dimensional machine room virtual model, a globally unique device identifier is assigned to each physical device to achieve accurate mapping and data association. The device identifier can be an IP address, a MAC (Media Access Control Address) address, an asset number (Asset Tag), or a unique ID generated by a configuration management database (CMDB). The choice of identifier depends on the existing operation and maintenance system specifications and the availability of data sources. For example, if the monitoring system is based on IP address for device management, then using the IP address as the unique identifier is the most direct choice. In the three-dimensional model creation or import stage, each device model needs to be bound to its actual corresponding unique identifier. This process is usually completed during the model data preparation stage and can be done by embedding the metadata in the model file or by assigning the device identifier to each graphical object dynamically through a program script after the three-dimensional engine loads the model. Therefore, one-key positioning can be achieved according to the target device identifier of the target monitored server, i.e., determining the server graphical object of the target monitored server in the three-dimensional machine room virtual model. To visually locate the faulty device, the display state of the server graphical object on the visualization page can be updated to a fault state. The device state can be distinguished by color, with green indicating normal state, yellow indicating warning state, and red indicating fault state. Of course, in addition to color differentiation, more rich visual styles can be defined for the device state, such as different light intensity, transparency, or metallicity, to further distinguish the state visually. To facilitate the maintenance personnel to view the fault reason, the fault data interaction tag of the server graphical object is generated according to the fault data information. The fault data interaction tag can be displayed to the user by clicking the server graphical object with the mouse. The fault data interaction tag contains the fault reason and can further include maintenance suggestions, concentrating on displaying the processed data. The interface is highly visual, and the fault information required can be easily obtained in multiple dimensions without extensive training.

[0017] In the technical scheme provided in the embodiment, the data acquisition terminal is used to acquire data of the monitored server, the acquired data is uploaded to the edge server for initial analysis to determine whether there is an exception, only the data with the exception is sent to the central server for fault identification, useless data transmission is prevented from occupying a large amount of bandwidth, on the basis of not occupying network resources, the delay caused by data transmission is reduced, the central server displays the fault on a visual page of the monitoring terminal after identifying the fault, the software and hardware information acquisition and information processing are integrated, through the progressive data processing flow, not only the hardware deployment cost is effectively reduced without deploying different products and different functions of equipment multiple times, but also the fault identification accuracy is improved, the fault is positioned and displayed on the monitoring terminal after identifying the fault, the operation and maintenance personnel can directly locate the fault equipment, and do not need to switch among multiple systems, the fault monitoring and fault repair efficiency is effectively improved, the fault response delay is maximally reduced, the whole process uses the automatic processing mode to reduce the business accidents caused by fault omission, and the stability and reliability of the server room are ensured.

[0018] The above embodiment does not make any limitation on the data acquisition terminal 1, and the present application also provides an exemplary implementation manner, which can include the following content. The data acquisition terminal 1 includes a hardware sensor group, a computer readable storage medium, a data processor, a scanner and a data interface; the data acquisition terminal 1 is connected with the corresponding edge server 2 through the data interface; the hardware sensor group includes temperature sensors, humidity sensors and vibration sensors arranged at air outlets of each monitored server in the server room, and the sensor acquisition data is taken as environment perception data and sent to the data processor; the scanner reads device label information of the monitored server and sends the device label information to the corresponding edge server 2 through the data interface; the computer readable storage medium stores a computer program, the computer program is injected into a target position of a computer program code of the monitored server, at least obtains running state data and various log data of the monitored server, and after adding a time stamp and a device identifier of the monitored server, sends the running state data and the various log data to the corresponding edge server 2; the data processor extracts temperature values, vibration values and humidity values from the environment perception data, obtains environment monitoring data of the monitored server, and after adding a time stamp and a device identifier of the monitored server, sends the environment monitoring data to the corresponding edge server 2 through the data interface.

[0019] In the embodiment, as Figure 3As shown, the hardware sensor group of the data acquisition terminal 1 can include a temperature probe (for monitoring temperature, whether high temperature occurs), a humidity sensor (for monitoring humidity, whether in the appropriate range), a three-axis vibration sensor (for monitoring machine vibration, whether abnormal vibration occurs). The data interface can be a fiber optic network port, compatible with USB3.0 data transmission. The scanner can scan two-dimensional code or bar code information, so that the device label information (including IP, serial number, maintenance record) can be automatically read. The hardware sensor group can collect basic data (temperature, humidity, current) according to a user-defined frequency, such as every 10 seconds, and when the detected parameter exceeds the threshold value, the sampling frequency is automatically increased to 2 times per second. After all data are added with time stamps and device numbers, they are uploaded to the edge server 2 through double channels (wired + backup wireless). The computer program of the computer readable storage medium can automatically inject the key positions (such as function entry / exit) of the monitored server software, realizing non-interrupted data acquisition. The data processor performs a filtering process on the raw data collected by the sensor, extracts specific data (temperature, humidity, vibration), and uploads the processed data to the edge server 2 through efficient data transmission to reduce response delay.

[0020] The above embodiments do not limit how to generate the abnormal report log and the selection mode of the to-be-identified data. Based on the above embodiments, the present application further provides an exemplary implementation mode, which can include the following contents: For the case that the abnormal log is a hardware alarm abnormal log, the target device identifier of the target monitored server, the abnormal alarm log, the abnormal detection data and the abnormal alarm time are extracted from the target server monitoring data, and the environment log is generated according to the target device identifier, the abnormal hardware identifier, the abnormal alarm type corresponding to the abnormal alarm log, the abnormal detection data and the hardware abnormal alarm time; the to-be-identified running state data of the target monitored server at the hardware abnormal alarm time, the target running state parameter exceeding the target running state parameter threshold value and the duration thereof are extracted from the target server monitoring data, and the device running log is generated according to the target running state parameter and the duration thereof, the target device identifier, the hardware abnormal alarm time and the software identifier; the abnormal report log is generated according to the environment log and the device running log, and the to-be-identified running state data and the abnormal alarm log are taken as the to-be-identified data.

[0021] For the case that the abnormal log is a software alarm abnormal log, the target device identifier of the target monitored server, the software abnormal alarm time, the software alarm abnormal log, the to-be-identified running state data at the software abnormal alarm time, the target running state parameter exceeding the target running state parameter threshold and the duration thereof are extracted from the target server monitoring data, and the device running log is generated according to the target running state parameter and the duration thereof, the target device identifier, the software abnormal alarm time, the software identifier and the software alarm type corresponding to the software alarm abnormal log; the target environment monitoring data corresponding to the software abnormal alarm time is extracted from the target server monitoring data, and the environment log is generated according to the target device identifier, the hardware identifier and the software abnormal alarm time; the abnormal report log is generated according to the environment log and the device running log, and the to-be-identified running state data, the software alarm abnormal log and the target environment monitoring data are taken as the to-be-identified data.

[0022] In the embodiment, for a large number of normal logs, only one non-exception prompt information is displayed after merging, and the information concerned in the abnormal log is extracted, for example, the time stamp, the device number, the module identifier, the key parameter, the important value, the event supplement (which can or can not be present), and these concerned information is generated into a short log to replace the log data. For example, the format of the environment log can be, for example, [2025-08-13T19:51:03.452] [A25-C35] [hardware HW] [temperature alarm TEMP ALERT] outlet 89.2 degrees (threshold 85 degrees). The format of the device running log can be, for example, [2025-08-13T19:51:03.755] [010-A01] [software SW] [CPU load] 98% 15s duration. The format of the abnormal report log can be, for example, [2025-08-13T19:51:04.128] [C6669] [HW / SW] [critical alarm CRITICAL] memory temperature 92 degrees + process crash (threshold 80 degrees + PID: 4412). If related data needs to be counted according to the fault type, a statistical summary log can also be generated, and the format of the statistical summary log can be, for example, [2025-08-13T19:51:10.000] [10 cabinet 1] [count COUNT] 10 minutes of abnormal count: temperature TEMP (temperature sensor) x 3 | memory MEM (memory) x 1 | network card NET (network card) x 0.

[0023] The above embodiments do not make any limitation on how to construct the three-dimensional machine room virtual model, and the application further provides an exemplary construction implementation of the three-dimensional machine room virtual model. Considering that the central server 3 usually adopts a high-performance server with high computing capability, the embodiment constructs the three-dimensional machine room virtual model through the central server 3, and then sends the constructed three-dimensional machine room virtual model to the monitoring terminal 4, which can include the following contents: The central server 3 acquires physical space data and equipment asset information of the server machine room, determines a model hierarchical structure for organizing each entity object of the server machine room, and generates unique identification information for each interactive entity object of the server machine room; based on the model hierarchical structure, three-dimensional modeling of the server machine room is performed according to the physical space data and the equipment asset information, and the three-dimensional model is rendered to obtain a three-dimensional model corresponding to the server machine room; wherein the target equipment graphical object of the three-dimensional model corresponding to the server machine room at least includes user data attributes, color configuration items and interactive labels; cabinet management information, case management information, mainboard management information, mainboard system information and hard disk information of the server machine room are acquired to generate equipment state data sets; based on the unique identification information of each interactive entity object, the equipment state data of each interactive entity object is acquired from the equipment state data sets, the corresponding equipment graphical object of each interactive entity object in the three-dimensional model corresponding to the server machine room is located, and the corresponding equipment state data is filled into the user data attributes, and the corresponding color configuration items are updated at the same time, to obtain a three-dimensional machine room virtual model, and the three-dimensional machine room virtual model is sent to the monitoring terminal 4.

[0024] In this embodiment, the physical space data of the three-dimensional machine room virtual model includes but is not limited to the architectural structure, size, internal layout, and location and specifications of all fixed facilities of the machine room, which can be obtained through the architectural design drawings of the machine room, such as CAD (Computer-Aided Design) format floor plan, elevation and section view. In actual operation, the operation and maintenance personnel or modelers need to carefully analyze these CAD drawings to extract key spatial information. By importing DWG or DXF files into three-dimensional modeling software or specialized conversion tools, the basic wall contour of the machine room can be automatically generated, greatly simplifying the repetitive work in the early stage of modeling. In addition to the architectural structure, floor information such as the material, size and laying method of the anti-static floor is also needed, which is crucial for simulating the real environment of the machine room. In some cases, if the CAD drawings are missing or do not match the reality, on-site investigation may be needed, using tools such as laser range finders for field measurements to ensure the accuracy of the data. After obtaining the physical space data of the machine room, detailed information of all IT equipment and related assets needs to be obtained, including servers, switches, routers, storage devices, UPS (Uninterruptible Power Supply), precision air conditioners, power distribution cabinets, surveillance cameras and all other key equipment deployed in the machine room. For each type of equipment, multi-dimensional information needs to be collected: physical attributes such as the brand, model, precise length, width and height of the equipment, and its specific installation location in the machine room (e.g. which U position of which cabinet), logical attributes such as the asset number, IP address, MAC address of the equipment, and the purpose of the equipment (such as web server, database server). This information can usually be obtained from the IT asset management system (CMDB) or network management system. Associating these physical and logical information is the basis for subsequent implementation of equipment state visualization and intelligent positioning functions. For example, during modeling, a unique identifier needs to be assigned to each equipment model, which can correspond to the equipment number or IP address in the asset management system, so as to realize precise control and data binding of specific equipment in the three-dimensional scene. In order to ensure the unity, maintainability and high performance of the three-dimensional machine room virtual model, the model structure, naming convention, material map, detail level and other aspects need to be determined. First, in terms of model structure, a clear hierarchical structure should be used to organize the objects in the scene. For example, the entire machine room can be taken as a top-level group, which contains architectural structures, cabinets, environmental equipment and other sub-groups, and each cabinet group contains server, switch and other equipment models. This hierarchical organization method is not only convenient for management, but also conducive to subsequent interactive operation and performance optimization. Second, in terms of naming convention, a unified and meaningful naming rule must be established for each interactive object in the scene (such as cabinet, server).For example, the format cabinet-001, server-192.168.1.10 can be used to ensure the uniqueness and readability of the name. Standard naming is crucial for achieving fast retrieval of devices and data binding. In terms of materials and textures, the specification should clearly specify the size, format (such as JPG, PNG), and UV mapping standards of the textures to ensure consistency and efficiency in rendering the model on the web. Finally, in terms of detail level, different precision versions of the model are created based on the importance of the device and the distance from the camera to balance visual effects and rendering performance.

[0025] Modeling the building structure of the machine room is the foundation of constructing the virtual scene. This process is usually done in professional 3D modeling software. First, the CAD floor plan of the machine room can be imported as a reference base map. Based on the base map, basic structures such as walls, floors, ceilings, doors, windows, and columns are accurately created using polygon modeling or spline extrusion techniques. For example, the two-dimensional outline of the wall can be drawn first, and then the height can be given through the extrusion modifier to form a three-dimensional wall. For doors and windows, the corresponding openings can be cut on the wall through Boolean operations, and then the pre-made door and window models can be installed. To improve efficiency and accuracy, array tools can be used to quickly generate repeated structural elements, such as the grid of anti-static floors. During the modeling process, strict adherence to the size specifications established earlier is necessary to ensure that the spatial proportions of the virtual machine room are completely consistent with the physical machine room. After completing the basic structure modeling, the model can be given appropriate materials, such as latex paint for walls and metal texture for floors, to enhance the realism of the scene.

[0026] Since three-dimensional models need to be rendered in real-time in web browsers, their complexity and data volume directly affect system performance and user experience. Therefore, model optimization and lightweight are necessary to reduce the number of polygons and the size of textures as much as possible while ensuring acceptable visual effects. The construction process of the three-dimensional model corresponding to the server room can include the following contents: Based on the fact that the spatial proportions of the virtual machine room are the same as those of the physical machine room, the architectural structure virtual model of the server room is generated according to the physical space data. According to the size of the three-dimensional machine room virtual model and the fault monitoring requirements, the first type of target device is modeled through geometric bodies, the second type of target device does not display geometric details, and the third type of target device replaces the geometric body model with a texture. According to the type of each device in the device asset information, the corresponding three-dimensional modeling of each device is performed, and based on the model hierarchy, each device is placed in the corresponding position of the architectural structure virtual model to obtain the three-dimensional model corresponding to the server room.

[0027] First, the principle of "just enough" should be followed in the modeling stage to avoid creating unnecessary geometric details. For example, for devices far from the camera, a low-polygon model can be used; for small or secondary devices, their structure can be appropriately simplified. Second, maps can be used to represent the details of the model. For example, the vents, buttons, labels, etc. on the cabinet can be completely realized by a high-precision transparent map, rather than by complex geometry, which can greatly reduce the number of faces. In addition, normal maps can be used to simulate the concave and convex details of the surface, improving the texture of the model without increasing the geometric complexity. Before exporting the model, the modeling software's optimization tools can be used to automatically reduce the number of polygons in the model. Finally, in terms of mapping, the size of the map should be reasonably controlled to avoid using excessively large texture files, and the map should be compressed to reduce memory usage and loading time.

[0028] After building the three-dimensional virtual model of the server room, data synchronization and state mapping between the physical world and the three-dimensional virtual model need to be implemented, that is, the real-time device state data of the server (such as running normally, warning, failure) is associated with the corresponding device graphical object in the three-dimensional virtual model, and is dynamically rendered through intuitive visual coding (color differentiation). In order to associate business data (such as device ID, IP address, state information) with device graphical objects (such as Mesh representing servers), the custom attribute mechanism provided by the three-dimensional engine can be used. Each device graphical object in this embodiment has a user data attribute, which is specifically used to store user-defined data. The user data attribute can be a normal JavaScript object, which will not be used by the internal logic of the three-dimensional engine, so it can safely store any application-related data without conflicting with the engine's own functions. Directly modifying the attributes of device graphical object instances is avoided, preventing potential naming conflicts and compatibility issues that may arise from future engine upgrades. By filling in the user data attribute, each device graphical object carries rich business information. When locating a device graphical object, you only need to traverse the device graphical objects in the scene graph and check the corresponding attribute value in its user data attribute. In order to improve operational efficiency and situational awareness, different device states (normal, warning, failure) are assigned clear and easily distinguishable colors, and the device state is dynamically rendered on the three-dimensional model in an intuitive color coding manner, so that the operator can quickly identify the devices that need attention in a complex room layout. In the implementation of this embodiment, these colors can be defined as global constants or configuration items in the application to facilitate unified management and modification. This definition makes the code more readable and maintainable. When the color needs to be adjusted, only this one place needs to be modified to take effect globally.

[0029] When the central server 3 generates the three-dimensional machine room virtual model, it can be sent to the monitoring terminal 4. The monitoring terminal 4 realizes real-time visualization of the device state, and can use polling or server push to obtain real-time data. The received device state data list is traversed, and for each device in the list, the unique identifier is used to find the corresponding model object in the three-dimensional virtual scene through the get Objects By Property (command) method. When the target object is found, the state information of the device (for example, "status":"warning") is parsed, and the visual representation of the model object is updated according to the preset rules. For example, the color attribute of the model material can be modified, or a material representing the warning state can be switched to. At the same time, the state field in the model user data attribute can also be updated to maintain data consistency.

[0030] The three-dimensional machine room virtual model constructed based on the above embodiment, the server graphic object state updating and fault data interaction tag generation process can include: The fault color value corresponding to the fault state is obtained, and the color value of the color configuration item of the server graphic object is modified to the fault color value. The fault device, fault device parameter and fault time are extracted from the fault data information, and the fault device, fault device parameter and fault time are filled into the interaction tag to obtain the fault data interaction tag.

[0031] Further, the present application also supports autonomous selection of display of all or part of the device all or custom information: the monitoring terminal 4, when receiving the running state display requirement request, obtains the to-be-displayed running state parameter and the display condition, such as displaying all devices with temperature higher than 40 degrees, displaying devices with storage usage rate exceeding 60%, if the to-be-displayed running state parameter includes the to-be-displayed device identifier, the corresponding target graphic object is determined in the three-dimensional machine room virtual model according to the to-be-displayed device identifier, and the target running state data corresponding to the to-be-displayed running state parameter is obtained from the user data attribute of the target graphic object. If the target running state data meets the display condition, the color configuration item of the target graphic object is modified according to the display condition; if the to-be-displayed running state parameter does not include the to-be-displayed device identifier, the user data attributes of each device graphic object of the three-dimensional machine room virtual model are traversed in turn, the running state data corresponding to the to-be-displayed running state parameter is obtained, and each target graphic object meeting the display condition is determined. The color configuration item of each target graphic object is modified according to the display condition.

[0032] As can be seen from the above, the present embodiment can easily obtain all information of the device in multiple dimensions through three-dimensional visualization display, and can centrally display and process the data. The interface has high visualization, which is conducive to improving the fault response efficiency.

[0033] Further, in order to reduce the fault response delay, based on the above embodiment, the application also provides dynamic adjustment of data compression rate and transmission priority according to network status, which can include the following contents: The edge server 2 acquires the first type of network parameter data of the communication link between itself and each data acquisition terminal 1 respectively, the second type of network parameter data of the communication link between itself and the central server 3; determines the first network status according to the first type of network parameter data, and determines the second network status according to the second type of network parameter data; matches the corresponding first data compression rate according to the first network status, and sends the first data compression rate to the data acquisition terminal 1, so as to compress the server monitoring data according to the first data compression rate; matches the corresponding second data compression rate according to the second network status, and compresses the to-be-recognized data according to the second data compression rate; if the current bandwidth utilization rate is lower than the preset threshold, the data with the highest transmission priority is preferentially transmitted according to the data type transmission priority, if the to-be-transmitted data includes multiple data types, the to-be-transmitted data is split into data segments according to the data type transmission priority, and the data segment with the highest transmission priority is preferentially transmitted; if the current network is interrupted, the current to-be-transmitted data is stored in the local target storage space.

[0034] Among them, the value of data compression rate is inversely proportional to network status, that is, the better the network status, the lower the data compression rate, the network parameter data can include bandwidth, delay and packet loss rate, the worse (low bandwidth, high delay) the network status, the higher the compression rate, on the contrary, the better the network status, the lower the compression rate. The preset bandwidth and delay threshold (such as bandwidth threshold: 10Mbps, 50Mbps; delay threshold: 50ms, 100ms), after comparing the real-time data of the current network status with the preset threshold, the matching compression rate (such as compression rate 90%, 70%) is determined, for example, when the bandwidth <10Mbps or the delay >100ms, the compression rate is more than 90% to reduce the data volume and preferentially transmit the critical alarm information; when the bandwidth >50Mbps and the delay <50ms, the compression rate is about 75%, which ensures the complete transmission of data. Combined with the actual use scene, for example, lossless compression algorithm can be used, and the volume is reduced by at least 75%. Further, the data type can also be preferentially transmitted according to the actual network status, such as alarm data preferentially occupying bandwidth and basic data allowing delay transmission. Further, it also supports offline buffering, that is, when the network is interrupted, 72 hours of data can be stored locally.

[0035] Further, based on the above embodiment, the application also provides a multi-dimensional alarm and alarm pushing implementation manner, as shown in Figure 3 The implementation manner can include the following contents: The central server 3 determines the fault emergency level, the fault equipment type, the fault frequency and the fault safety level according to the fault data information, generates multi-dimensional alarm information according to the fault emergency level, the fault equipment type, the fault frequency and the fault safety level, matches the authorized user terminal in the alarm information pushing database according to the multi-dimensional alarm information, and pushes the multi-dimensional alarm information to the authorized user terminal, matches the alarm display terminal and the alarm prompt terminal in the alarm information pushing database according to the multi-dimensional alarm information, and controls the multi-dimensional alarm information to be displayed on the visual page of the alarm display terminal and to be prompted on the alarm prompt terminal.

[0036] Among them, the emergency alarm level, the important alarm level, the secondary alarm level and the prompt alarm level can be set according to the emergency degree. The emergency alarm level is the risk of affecting the core service or important data of the system, such as server downtime causing complete business interruption, core database crash causing access failure, which needs to be handled immediately. The important alarm level is the service quality decline, which needs to be repaired urgently, such as network bandwidth overload causing service delay surge, CPU (Central Processing Unit) usage rate continuously exceeding 90% affecting business performance, server high temperature causing overall performance reduction. The secondary alarm level is not yet affecting the business, but needs to be handled later, such as disk remaining space less than 10% needing to be cleaned up, non-critical service running abnormally. The prompt alarm level is the potential risk or state prompt, such as CPU usage rate approaching 80% threshold needing to be monitored, log file size exceeding expected growth rate. According to the device type, hardware class alarms can be set, including power failure, CPU / memory / hard disk abnormality, fan speed abnormality and other hardware component problems, such as hard disk failure or unstable power supply voltage. System class alarms: related to operating system or software abnormalities, such as system crash, service unresponsive, resource shortage. Network class alarms: covering all network related problems, such as connection interruption, high delay, abnormal packet loss rate. Security class alarms: for potential security threats, such as virus attack, abnormal login, distributed denial of service attack. Storage class alarms: all storage related components have problems, such as insufficient disk space, redundant disk array card failure, file system error, etc. Environment class alarms: related to physical environment, such as server temperature too high, abnormal cooling, etc. According to the occurrence frequency, high frequency alarms can be set: repeated triggering in a short time, which may be caused by system design defects or long-term resource shortage, such as CPU / memory usage rate continuously exceeding threshold, network interface frequent packet loss or delay jitter. Medium frequency alarms: triggered at fixed intervals (such as daily / weekly), which may be related to business peak or timing task, such as database connection pool depletion during daily business peak, timing backup task failure due to some reason. Low frequency alarms: low probability of occurrence, but may imply potential risks, such as security scan finding unpatched vulnerabilities, server occasional restart. Burst alarms: a large number of alarms occur in a short time, usually caused by sudden events, such as virus intrusion causing resource occupation surge, database lock performance causing a large number of transaction blocking. Intermittent alarms: irregular and difficult to reproduce, such as random error codes in logs. According to the secret level, top secret level alarms can be set: high sensitive alarms related to system core security or business core data, such as any alarm in core machine room, high sensitive data machine room server intrusion alarm. Confidential level alarm: alarm affecting business continuity or containing sensitive operation data, leakage may cause commercial competition risk, such as customer data server storage alarm.Secret level alarm: routine operation type alarm, which may leak the system stability but has no direct security threat, such as server resource overload, service response delay alarm.

[0037] When the corresponding alarm information is generated according to the multi-dimensional alarm generation rule, the corresponding push can be combined with the user permission, the alarm information push database is pre-set, the alarm information push database includes the permission management of different users, the position, rank and permission of each user are set, so that according to the database, the push can be made to the user with corresponding permission according to the emergency degree, device type, frequency of occurrence, secret level, such as the problem of core data room can only be pushed to the user with core permission, and the problem of memory only needs to be pushed to the hardware interface. In order to further improve the effectiveness of the alarm, the embodiment adopts a plurality of information push modes, for example, the AB angle redundancy mode can be used to ensure that important alarms can be handled in time. The authorized user end can be a mobile APP, which can push classified alarm notifications such as top secret, secret, emergency and important (customized) in real time, support to view detailed alarm information and fault information, support to view three-dimensional positioning view by clicking, and support to call camera permission.

[0038] Further, the monitoring terminal 4 can also set a web console, which supports to display the room information according to the permission, supports to filter the equipment according to multiple conditions, updates all information in real time, and supports to customize the displayed information. The alarm prompt end can include a voice end and a mobile end, and for classified alarms such as top secret, secret, emergency and important (customized), short message notification and voice call reminder can be supported at the same time. Further, the visual page of the monitoring terminal 4 also includes a device detail panel and an alarm processing board, the device detail panel displays customized device location, IP and other information, and clicking any device can view real-time parameter curve (temperature, load, network traffic), maintenance history record and associated device topology. The alarm processing board can sort the pending tasks according to the emergency degree, preferentially prompt the fatal alarm and perform the top highlight reminder, and the alarm which has not been processed for a long time will be repeatedly pushed and reminded; the standard disposal process guide is provided, including processing suggestion and historical processing process, and fast troubleshooting method is provided; the alarm to-be-done work order is generated in real time, reminding the administrator to process as soon as possible, and the work order can be ended after processing, so as to realize fast response and rapid processing.

[0039] As can be seen from the above, the embodiment can analyze the alarm push from multiple dimensions through permission management and accurate push, accurately and in multiple ways to alarm, which not only can play the effectiveness of the alarm, but also can prevent data leakage and effectively improve the security.

[0040] In order for those skilled in the art to clearly understand the technical solutions of the present application, the present application also provides an exemplary implementation, which can include the following contents: For the scenario of hard disk fault handling, when the data collection terminal 1 detects abnormal vibration of the target monitored server hard disk through the vibration sensor, exceeding the set threshold. The server fault monitoring system automatically associates the log of the device at the same time, and finds that the number of remapping sectors exceeds the threshold. The push layer sends an alarm to the mobile APP and web console at the same time, including the device location (three-dimensional machine room virtual model positioning), the recommended inspection item (hard disk module), the associated device list (other nodes in the same storage cluster), the historical process of handling the fault and the suggestion (such as replacing the hard disk), and the push object is the storage interface. The operation and maintenance personnel replace the hard disk according to the prompt, and the system automatically updates the maintenance record and removes the alarm.

[0041] For the software fault scenario, the central server identifies that the abnormal report log includes the CE keyword, determines that the number of memory CE alarms of a certain memory in the core data room within a short time exceeds the preset number threshold through the to-be-identified data, and determines that the server device identification to which the memory belongs is C-A-1-54688. Push the abnormal information to the administrator, including alarm information, location information, memory information, historical processing process and suggestion, generate a to-do work order in real time, and the push object is the hardware interface of the core data. The machine room administrator immediately goes to the machine room for maintenance with the same memory, and removes the alarm and ends the to-do work order after the maintenance is completed.

[0042] Based on the above embodiment, the central server 3 will receive multi-source data. In order to improve the data processing efficiency of the central server, the embodiment also provides an implementation manner of multi-source data fusion analysis, which can include the following contents: First, align the device physical parameters (temperature, vibration), network status (IP, traffic), and software log (error code) on the time axis. Then, calibrate the multi-dimensional time data through the dynamic time warping algorithm, and automatically associate each log data through the sliding time window.

[0043] Wherein, according to the scene to determine the protocol adopted by the global time synchronization error requirements, error budget < 1ms, the central server 3, hardware sensor group, edge server 2, monitoring terminal 4 network unified PTP protocol (Precision Time Protocol, precision time protocol) is started. If the error budget is < 10ms, NTP (Network Time Protocol, network time protocol) v4 can be used to maintain time synchronization. For sensors, edge servers 2, central servers 3, three types of nodes, unified according to the "four tuple" encapsulation data: log: event ID, level, text, local timestamp. Device information: model, firmware version, MAC, location coordinates. Sensor data: sample value, dimension, unit, sampling period. Network information: egress bandwidth, RTT, packet loss rate, link layer type. Further, a 64-bit "synchronization time stamp" field can be added to the four tuple header, with a precision of 100ns. Each sensor can send a data frame to the edge server every 20ms through the UDP (User Datagram Protocol, User Datagram Protocol) / DTLS (Datagram Transport Layer Security, Datagram Transport Layer Security) tunnel, and the edge server can send a batch of compressed packages to the central server every 50ms through the HTTP2 (HyperText Transfer Protocol version 2, HyperText Transfer Protocol version 2) long connection. The edge server can pre-set a local cache with a memory ring queue, and the queue overflow triggers the local DTW (Dynamic Time Warping, Dynamic Time Warping) pre-alignment, only the "calibrated timestamp" is pushed to the upstream, reducing the load of the central server. The central server takes the central server local PTP domestic clock as the absolute time axis, and for log type data, the "event level + keyword hash" is extracted. For sensor data, the "sample value first difference" can be taken. For network data, the "RTT (Round-Trip Time, Round-Trip Time) jitter" can be taken. For device information, the "firmware version number" can be taken to identify the mutation caused by hot upgrade. Set the window length to 2s and the step to 500ms, and do z-score normalization for each sequence, then use Sakoe-Chiba (band constraint strategy) constraint (bandwidth 100ms) to find the shortest regular path, output "time shift Δt" and "frequency drift slope k", precision 10µs. Map the original timestamp to the global axis through Δt, k, and generate "unified timestamp". Write the correction record back to the "time alignment table" (TTL (Time To Live, Time To Live) 24h) for data traceability tracking.

[0044] For 500ms sliding window associated log operation: the length of the window is predefined as 500ms, and the step is 100ms, i.e. the window overlaps 400ms, and the left boundary of the window = current global time - 500ms. The following association rules are set according to the priority from high to low: ① time: the unified timestamp falls within the window; ② event: the log event ID is consistent with the sensor trigger flag; ③ device: the same device ID or MAC is used; ④ network: the same link layer type and RTT difference < 5ms. If multiple logs in the same 500ms window point to the same event, the one with the "unified timestamp" closest to the midpoint of the window is selected. If the log is missing, but the sensor and network information indicate the event, the log field is completed by interpolation, and can be marked as "trusted level = low".

[0045] As can be seen from the above, by highly unifying and integrating all data, the embodiment is advantageous for quickly identifying faults, shortening fault processing time, and effectively reducing the possibility of false positives.

[0046] The skilled person can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0047] The above has carried on the detailed introduction to the server fault monitoring system provided by the present application. The principle and implementation mode of the present application are described by applying specific examples in this paper, and the above example description is only applicable to help understand the method and core idea of the present application. It should be pointed out that for ordinary skilled person in the technical field, some improvements and modifications can be made to the present application without departing from the principle of the present application, and these improvements and modifications also fall within the protection scope of the present application.

Claims

1. A server fault monitoring system, characterized in that: It includes data acquisition terminals deployed in the server room, edge servers connected to the corresponding regional data acquisition terminals, a central server connected to each edge server, and a monitoring terminal connected to the central server; Each data collection terminal generates server monitoring data based on the device identification, environmental monitoring data, operating status monitoring data, log data and data collection time of the monitored server and sends it to the connected edge server; Each edge server, when there is an abnormality log in the target server monitoring data, extracts the data to be identified related to the abnormality log from the target server monitoring data, and generates an abnormality reporting log for the target monitored server; The central server, when determining that a target monitored server has a fault based on the exception reporting log and the data to be identified, generates fault data information and sends it to the monitoring terminal; The monitoring terminal determines the server graphic object of the target monitored server in the three-dimensional computer room virtual model based on the target device identifier of the target monitored server, updates the display state of the server graphic object on the visualization page to a fault state, and generates a fault data interaction tag for the server graphic object based on the fault data information.

2. The server fault monitoring system according to claim 1, characterized in that: Before determining the server graphic object of the target monitored server in the three-dimensional computer room virtual model according to the target device identifier of the target monitored server, the method further includes: The central server obtains physical space data and equipment asset information of the server room, determines a model hierarchical structure for organizing various entity objects in the server room, and generates unique identification information for each interactive entity object in the server room; Based on the model hierarchical structure, three-dimensional modeling of the server room is performed according to the physical space data and the equipment asset information, and the three-dimensional model is rendered to obtain a three-dimensional model corresponding to the server room; wherein the target device graphic object of the three-dimensional model corresponding to the server room includes at least user data attributes, color configuration items, and interactive labels; Obtain the cabinet management information, chassis management information, motherboard management information, motherboard system information, and hard disk information of the server room to generate a device status data set; Based on the unique identification information of each interactive entity object, the device status data of each interactive entity object is obtained from the device status data set, the device graphic object corresponding to each interactive entity object is located in the three-dimensional model corresponding to the server room, and the corresponding device status data is filled into the user data attribute. At the same time, the corresponding color configuration items are updated accordingly to obtain a three-dimensional virtual model of the server room, and the three-dimensional virtual model of the server room is sent to the monitoring terminal.

3. The server fault monitoring system according to claim 2, characterized in that: Updating the display state of the server graphic object on the visualization page to a fault state, and generating a fault data interaction tag of the server graphic object according to the fault data information, including: Obtaining a fault color value corresponding to a fault state, and modifying the color value of the color configuration item of the server graphic object to the fault color value; The faulty device, faulty equipment parameters and fault time are extracted from the fault data information, and the faulty device, faulty equipment parameters and fault time are filled into the interactive tag to obtain a fault data interactive tag.

4. The server fault monitoring system according to claim 2, characterized in that: Based on the model hierarchical structure, three-dimensional modeling of the server room is performed according to the physical space data and the equipment asset information, including: Based on the fact that the spatial proportion of the virtual server room is the same as that of the physical server room, generating a virtual model of the architectural structure of the server room according to the physical space data; According to the scale of the three-dimensional computer room virtual model and the fault monitoring requirements, determining a first category of target devices modeled by geometric bodies, a second category of target devices without displaying geometric details, and a third category of target devices whose geometric bodies are replaced by textures; According to the device type of each device in the device asset information, each device is modeled accordingly in three dimensions, and each device is placed in a corresponding position of the building structure virtual model based on the model hierarchy to obtain a three-dimensional model corresponding to the server room.

5. The server fault monitoring system according to claim 2, characterized in that: The monitoring terminal, upon receiving a request for displaying the operating status, obtains operating status parameters to be displayed and display conditions; If the operating status parameter to be displayed includes a device identifier to be displayed, determining a target graphic object corresponding to the device identifier to be displayed in the three-dimensional computer room virtual model, and obtaining target operating status data corresponding to the operating status parameter to be displayed from a user data attribute of the target graphic object; if the target operating status data satisfies the display condition, modifying a color configuration item of the target graphic object according to the display condition; If the operating status parameters to be displayed do not include the device identifier to be displayed, the user data attributes of each device graphic object of the three-dimensional computer room virtual model are traversed in sequence to obtain the operating status data corresponding to the operating status parameters to be displayed, and the target graphic objects that meet the display conditions are determined, and the color configuration items of each target graphic object are modified accordingly according to the display conditions.

6. The server fault monitoring system according to any one of claims 1 to 5, characterized in that: The data acquisition terminal includes a hardware sensor group, a computer-readable storage medium, a data processor, a scanner and a data interface; The data acquisition terminal is connected to the corresponding edge server through the data interface; The hardware sensor group includes temperature sensors, humidity sensors and vibration sensors deployed at the air outlets of each monitored server in the server room, and sends the sensor collected data as environmental perception data to the data processor; The scanner reads the device tag information of the monitored server and sends the device tag information to the corresponding edge server through the data interface; The computer-readable storage medium stores a computer program, which is injected into a target location of a computer program code of a monitored server, obtains at least operating status data and various log data of the monitored server, and sends the data to a corresponding edge server after adding a timestamp and a device identifier of the monitored server; The data processor extracts temperature, vibration and humidity values ​​from the environmental sensing data to obtain environmental monitoring data of the monitored server, and sends the data to the corresponding edge server through the data interface after adding a timestamp and a device identifier of the monitored server.

7. The server fault monitoring system according to any one of claims 1 to 5, characterized in that: The edge server obtains first-category network parameter data of the communication link between the edge server and each data acquisition terminal, and second-category network parameter data of the communication link between the edge server and the central server; determining a first network state according to the first type of network parameter data, and determining a second network state according to the second type of network parameter data; Matching a first data compression ratio corresponding to the first network state, and sending the first data compression ratio to the data acquisition terminal to compress the server monitoring data according to the first data compression ratio; matching a second data compression ratio corresponding to the second network state, and compressing the to-be-identified data according to the second data compression ratio; wherein the value of the data compression ratio is inversely proportional to the network state; If the current bandwidth utilization is lower than the preset threshold, the data with the highest priority will be transmitted first according to the data type transmission priority. If the data to be transmitted includes multiple data types, the data to be transmitted will be split into multiple data segments according to the data type transmission priority, and the data segment with the highest priority will be transmitted first. If the current network is interrupted, the data to be transferred will be stored in the local target storage space.

8. The server fault monitoring system according to any one of claims 1 to 5, characterized in that: The exception log is a hardware alarm exception log, which generates an exception reporting log of the target monitored server, including: Extracting the target device identifier, hardware alarm exception log, exception detection data, and hardware exception alarm time of the target monitored server from the target server monitoring data, and generating an environment log based on the target device identifier, hardware identifier, the exception alarm type corresponding to the hardware alarm exception log, the exception detection data, and the hardware exception alarm time; Extracting, from the target server monitoring data, the target monitored server's operating status data to be identified at the hardware anomaly alarm time, the target operating status parameter exceeding the target operating status parameter threshold and its duration, and generating a device operation log based on the target operating status parameter and its duration, the target device identifier, the hardware anomaly alarm time, and the software identifier; An abnormality reporting log is generated according to the environment log and the equipment operation log, and the operation status data to be identified and the hardware alarm abnormality log are used as data to be identified.

9. The server fault monitoring system according to any one of claims 1 to 5, characterized in that: The exception log is a software alarm exception log, which generates an exception reporting log of the target monitored server and includes: Extracting, from the target server monitoring data, the target device identifier of the target monitored server, the software abnormality alarm time, the software alarm abnormality log, the operating state data to be identified within the software abnormality alarm time, the target operating state parameter exceeding the target operating state parameter threshold and its duration, and generating a device operation log according to the target operating state parameter and its duration, the target device identifier, the software abnormality alarm time, the software identifier, and the software alarm type corresponding to the software alarm abnormality log; Extracting target environment monitoring data corresponding to the software abnormality alarm time from the target server monitoring data, and generating an environment log according to the target device identifier, hardware identifier, and the software abnormality alarm time; An abnormality reporting log is generated according to the environmental log and the equipment operation log, and the operation status data to be identified, the software alarm abnormality log, and the target environment monitoring data are used as data to be identified.

10. The server fault monitoring system according to any one of claims 1 to 5, characterized in that: The central server determines the fault emergency level, fault device type, fault frequency, and fault safety level based on the fault data information, and generates multi-dimensional alarm information; According to the multi-dimensional alarm information, matching the authorized user terminal in the alarm information push database, and pushing the multi-dimensional alarm information to the authorized user terminal; According to the multi-dimensional alarm information, the alarm display terminal and the alarm prompt terminal are matched in the alarm information push database, and the multi-dimensional alarm information is controlled to be displayed on the visualization page of the alarm display terminal and the alarm prompt is performed on the alarm prompt terminal.

Citation Information

Patent Citations

  • Three-dimensional machine room monitoring system and method

    CN102608939A

  • Server monitoring method and device, computer equipment and storage medium

    CN109815093A

  • Integrated data information monitoring platform and monitoring system

    CN110398927A

  • Cloud server state monitoring system and method

    CN111193643A

  • Data center intelligent operation and maintenance management system and method

    CN112269673A

Cited By

  • Server state monitoring method and electronic equipment

    CN121050971A

  • Server alarm function detection method and electronic equipment

    CN121396851A

  • IDC machine room fault point prediction method and system based on artificial intelligence and Internet of Things

    CN121786692A

  • IDC machine room fault point prediction method and system based on artificial intelligence and internet of things

    CN121786692B