A data center operation and maintenance visualization monitoring method combining big data and network topology
By combining big data and network topology with a data center operation and maintenance visualization monitoring method, the problem of false information misjudgment in data center operation and maintenance is solved, the accurate positioning and timely maintenance of equipment physical machines are achieved, and the operation and maintenance efficiency and stability are improved.
Patent Information
- Application Number
- CN202510943172.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-09
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-07-09
AI Technical Summary
Existing technologies make it difficult to avoid the impact of false information caused by data transmission interference and network attacks in data center operations and maintenance, which leads to misjudgment of the physical operation status of equipment. Traditional monitoring is also inefficient and difficult to expand.
A data center operation and maintenance visualization monitoring method that combines big data and network topology is adopted. By obtaining the operating status of the physical machine of the equipment, the U-bit ratio and the network transmission volume, the abnormality assessment value is calculated, and visual monitoring is performed in combination with the network topology structure to identify potential abnormal equipment and perform alarm maintenance.
It improves the scalability and accuracy of data center operation and maintenance, reduces the risk of misjudgment and missed judgment, and improves operation and maintenance efficiency and stability.
Smart Images

Figure CN120455321B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data center management technology, and specifically to a data center operation and maintenance visualization monitoring method combining big data and network topology. Background Art
[0002] With the rapid development of information technology, data centers have become critical infrastructure for enterprise operations and the digital transformation of society. As data centers continue to expand, the number of physical machines within them is growing exponentially, leading to increasingly complex system architectures. Furthermore, these machines are often provided by different vendors and run diverse operating systems and applications, placing extremely stringent demands on data center operational stability. Traditional data center operations and maintenance rely on inspections and simple command-line tools to obtain device status information, which is inefficient and prone to missing potential issues. Consequently, an intelligent and visual data center operations and maintenance approach is needed.
[0003] At present, some research is being conducted to improve the operation and maintenance efficiency of data centers. However, in existing technologies, while ensuring that data center operation and maintenance has high scalability, it is difficult to avoid being affected by data transmission interference and false information from network attacks, which may lead to misjudgment of the operation status of the physical equipment. For example, Jiao Hongbin's paper "Research on Intelligent Monitoring and Self-Healing Mechanism in Distributed Data Center IT Operation and Maintenance Management System" proposes a composition of distributed data center intelligent monitoring and anomaly detection of data center events to improve the operation and maintenance efficiency and operational stability of the data center. However, it is completely dependent on data collection from the data center and is sensitive to data noise and network attacks.
[0004] In order to solve the above technical problems, this application provides a data center operation and maintenance visualization monitoring method that combines big data and network topology to solve existing problems.
[0005] The data center operation and maintenance visualization monitoring method of this application that combines big data and network topology adopts the following technical solutions:
[0006] One embodiment of the present application provides a data center operation and maintenance visualization monitoring method combining big data and network topology, the method comprising the following steps:
[0007] For all physical machines in all cabinets in the data center, obtain the operating status of each physical machine at each moment, including temperature, memory usage, CPU usage, and network transmission volume; obtain the U-slot ratio of each cabinet;
[0008] The difference between each operating state of each device physical machine at each moment and the previous moment is used as the deviation value of each operating state of each device physical machine at each moment; the transient load of each device physical machine at each moment is calculated based on the overall distribution characteristics of the deviation values of all operating states of each device physical machine at each moment; based on the distribution range of the deviation values of all operating states of each device physical machine at each moment, combined with the U-bit proportion and the transient load, the abnormal assessment value of each device physical machine at each moment is calculated;
[0009] Perform visual monitoring of equipment physical machines based on abnormal assessment values, and issue alarms and repairs on faulty equipment physical machines.
[0010] In one embodiment, the process of obtaining the U-space ratio of each cabinet is as follows:
[0011] The U-bit capacity occupied by all physical devices in the cabinet is obtained by the number of physical devices on the cabinet; the U-bit capacity of the cabinet is determined by the container size of the cabinet; and the ratio of the U-bit capacity occupied by all physical devices to the U-bit capacity of the cabinet is used as the U-bit ratio of the cabinet.
[0012] In one embodiment, the process of obtaining the deviation value of each operating state of each physical machine of the device at each moment is as follows:
[0013] The absolute value of the difference between any item of operating status data of each device physical machine at the current moment and the previous moment is calculated, and recorded as the deviation value of any item of operating status data of each device at the current moment.
[0014] In one embodiment, the process of obtaining the transient load is as follows:
[0015] The sum of the deviation values of all items of operation status data of any device physical machine at the current moment is calculated as the transient load of the any device physical machine at the current moment.
[0016] In one embodiment, the process of obtaining the abnormality assessment value is as follows:
[0017] Calculate the range of the deviation values of all items of operating status data of each device physical machine at the current moment, and determine the abnormal assessment value of each device physical machine at each moment based on the range of the deviation values, the U-bit proportion and the transient load.
[0018] In one embodiment, the expression of the abnormality evaluation value of each physical machine of the device at each time is:
[0019] , where Indicates the abnormal evaluation value of the physical machine of the i-th device at the current moment, Indicates the instantaneous load of the physical machine of the i-th device at the current moment, Indicates the range of the deviation value of the physical machine of the i-th device at the current moment, Indicates the U-slot ratio of the cabinet where the physical machine of the i-th device is located.
[0020] In one embodiment, the visual monitoring of the physical machine of the device based on the abnormal evaluation value is specifically performed as follows:
[0021] Based on the abnormality assessment value, the potential abnormal device is obtained, and based on the continuous occurrence of the potential abnormal device over time, the physical machine of the faulty device is obtained; the potential abnormal device and the physical machine of the faulty device are visually marked.
[0022] In one embodiment, the process of obtaining the potential abnormal device is as follows:
[0023] The abnormal evaluation values of all device physical machines at any moment and all previous moments are used as the input of the cross-validation method, the output optimal threshold is used as the screening threshold at any moment, and the device physical machine whose abnormal evaluation value at any moment is greater than the screening threshold is marked as a potential abnormal device at any moment.
[0024] In one embodiment, the process of obtaining the physical machine of the faulty device is as follows:
[0025] If each device physical machine is marked as a potential abnormal device at a preset number of consecutive moments, each device physical machine is marked as a faulty device physical machine.
[0026] In one embodiment, the visual marking of the physical machines of potential abnormal devices and faulty devices is specifically as follows:
[0027] The intelligent operation and maintenance monitoring platform is connected to the data center through a network line. A visualization interface is reserved on the intelligent operation and maintenance platform. If the physical machines of each device in the data center are marked as potentially abnormal devices, the corresponding physical machines of the device in the visualization interface will be marked yellow. If the physical machines of each device are marked as faulty physical machines, the corresponding physical machines of the device in the visualization interface will be marked red.
[0028] This application has at least the following beneficial effects:
[0029] This application can expand the distributed data center network structure through the network transmission protocol, and improve the scalability of data center operation and maintenance; at the same time, compared with the traditional operation and maintenance monitoring process that simply judges the abnormal operation of the equipment physical machine through sensors, this application constructs an abnormal evaluation value for each equipment physical machine based on the correlation between the operation status data of each equipment physical machine in the data center and the U-bit ratio of the equipment physical machine, and evaluates the authenticity of the abnormality of the equipment physical machine; determines the faulty equipment physical machine based on the abnormal evaluation value, and realizes the accurate positioning and timely operation of the abnormal equipment physical machine; avoids the traditional monitoring process in which the operation status of the equipment physical machine is directly judged by the results of collecting the operation status data of the equipment physical machine in the data center, which is easily affected by data transmission interference and false information of network attacks, resulting in misjudgment of the operation status of the equipment physical machine, improves the operation and maintenance efficiency of the traditional data center operation and maintenance method, reduces the risk of misjudgment and missed judgment, and improves the stability and reliability of the data center operation. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0031] Figure 1 Flowchart of the data center operation and maintenance visualization monitoring method combining big data and network topology provided in this application;
[0032] Figure 2 This is a schematic diagram of the process of obtaining the U-space ratio of each cabinet. DETAILED DESCRIPTION
[0033] In order to further illustrate the technical means and effects adopted by this application to achieve the predetermined invention objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features and effects of the data center operation and maintenance visualization monitoring method based on the combination of big data and network topology proposed in this application. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics of one or more embodiments may be combined in any suitable form.
[0034] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.
[0035] The following describes in detail the specific scheme of the data center operation and maintenance visualization monitoring method combining big data and network topology provided by this application with reference to the accompanying drawings.
[0036] An embodiment of the present application provides a data center operation and maintenance visualization monitoring method that combines big data and network topology.
[0037] Specifically, the following data center operation and maintenance visualization monitoring method combining big data and network topology is provided. Figure 1 , the method comprises the following steps:
[0038] Step S1: For all physical machines on all cabinets in the data center, obtain the operating status of each physical machine at each moment, including temperature, memory usage, CPU usage, and network transmission volume; and obtain the U-bit ratio of each cabinet.
[0039] Traditional monitoring methods require installing a monitoring plug-in, such as Zabbix, on each physical device to monitor in the data center. However, in practice, the placement and removal of physical devices within each data center computer room may need to be adjusted to meet business needs. In this case, plug-in-based monitoring methods are not conducive to dynamic monitoring of physical devices.
[0040] Therefore, in order to automatically obtain the status of the physical machines in the data center, the network topology between all the physical machines in the data center is constructed as follows:
[0041] The intelligent operation and maintenance monitoring platform is connected to the data center via a network connection. Using an administrator account and password to log in and gain administrative privileges, the platform generates the data center's network topology by querying the network protocols of each physical device within the data center, traversing the IP addresses and MAC addresses of each physical device in each computer room and cabinet within the computer room. These network protocols include ARP tables, routing tables, and MAC tables.
[0042] For example, the monitoring platform links to obtain the routing table of the switch equipment physical machine in a single computer room, and can obtain the IP addresses of the equipment physical machines in different network segments within the current computer room local area network, as well as the IP addresses of the equipment physical machines in each network segment. Based on the comparison of the IP addresses and MAC addresses of the equipment physical machines, the cabinets to which each equipment physical machine in the computer room belongs and the connection status of the equipment physical machines are obtained, thereby forming the network structure of the equipment physical machines in the computer room. At the same time, based on the IP addresses of other computer room switches in the external network in the current computer room routing table, the network structure of the equipment physical machines in other computer rooms is obtained in the same way, thereby automatically generating the network topology of the current data center. Among them, the generation of the network topology structure of the equipment physical machines in the data center is a well-known technology, and the specific process will not be repeated here.
[0043] It should be noted that, under normal circumstances, the network connection structure of the data center does not change much in a short period of time. Therefore, this application performs network structure topology on a daily basis. At the same time, when the network connection of the physical machine of the data center equipment changes, it can also actively report the changes, and the network structure topology is also performed at this time.
[0044] In this way, the network topology between all physical machines in the data center can be obtained dynamically and in real time, avoiding the problem of poor scalability in traditional monitoring methods.
[0045] Furthermore, through the network topology of all physical machines in the data center, the operating status of the physical machines can be directly obtained through IP address access, such as querying through the ping command.
[0046] To reduce workload complexity, the intelligent monitoring and maintenance platform uses code to poll the operating status of all physical devices in the data center at fixed intervals. This allows users to obtain various operational status information, including temperature, memory usage, CPU utilization, and network traffic. Network traffic is measured by collecting statistics from network interfaces, capturing the number of data packets sent and received at fixed intervals.
[0047] In order to reduce the impact of data units and numerical values, the collected values of temperature and network transmission volume are normalized.
[0048] On the intelligent monitoring and operation and maintenance platform, the smaller the fixed time interval is, the higher the accuracy of data collection is. The general selection interval is [0.1s, 5s]. Preferably, in the embodiment of the present application, the fixed polling time interval is set to 1s.
[0049] At the same time, in order to obtain the U-bit resource status of each cabinet, a tag reader is deployed on each cabinet to obtain the U-bit electronic tag ID on each cabinet. In this way, the number of physical machines of equipment on the cabinet can be obtained through the U-bit detection interface, thereby obtaining the U-bit capacity occupied by all physical machines of equipment in the cabinet; further, since the container size of each cabinet is known, the U-bit capacity of each cabinet can be determined; further, the ratio of the U-bit capacity occupied by all physical machines of equipment in each cabinet to the U-bit capacity of the cabinet is calculated to obtain the U-bit share of the cabinet, and the U-bit share of the cabinet is reported to the intelligent monitoring and operation and maintenance platform.
[0050] Step S2, taking the difference between each operating state of each device physical machine at each moment and the previous moment as the deviation value of each operating state of each device physical machine at each moment; calculating the transient load of each device physical machine at each moment based on the overall distribution characteristics of the deviation values of all operating states of each device physical machine at each moment; calculating the abnormal evaluation value of each device physical machine at each moment based on the distribution range of the deviation values of all operating states of each device physical machine at each moment, combined with the U-bit proportion and the transient load.
[0051] Data center operations require real-time monitoring of the operating status of each physical machine. This ensures timely detection and rapid response to any abnormalities. Traditional data center operations rely on collected data on the operating status of physical machines. Safety thresholds are set for each piece of data. When the operating status of a physical machine reaches or exceeds the safety threshold, the machine is deemed to be operating abnormally. This safety threshold relies on empirical judgment and historical data analysis.
[0052] However, traditional O&M methods issue security alerts for each physical machine's operating status data, ignoring the interconnections between data. This makes it easy to misjudge network attacks and fall into attack traps. For example, during traditional O&M adjustments, the operating status of individual physical machines is analyzed. When a single physical machine experiences an abnormality, self-healing mechanisms are implemented for O&M management, such as remotely restarting the physical machine. However, if the abnormality data is forged by an attacker, the attacker can exploit the time it takes to steal critical data from the physical machine. Therefore, during the O&M process, it is necessary to combine the actual operating conditions of the physical machine to determine and perform O&M.
[0053] For example, a single physical machine in a data center might be responsible for a service or storage task within the data center. External access to this service is uncertain, so the machine's operational status data changes dynamically in real time, and this dynamic range varies within the machine's operating range. When an anomaly occurs, it often manifests as a significant instantaneous load change, approaching full load, causing the machine to operate at high intensity and potentially leading to a failure.
[0054] Based on the above analysis, the instantaneous load of each physical machine of the device at the current moment is calculated as follows:
[0055] For each item of operating status data of any device physical machine, calculate the absolute value of the difference between the current moment and the previous moment of the operating status data, and record it as the deviation value of the operating status data at the current moment; calculate the sum of the deviation values of all items of operating status data of any device physical machine at the current moment, and record it as the transient load of any device physical machine at the current moment.
[0056] It should be noted that in order to ensure the stability of data collection, data within the initial collection of 1 minute are not analyzed.
[0057] If a single physical machine has a large instantaneous operating load when the physical machine is running, the running status data of the physical machine increases significantly compared to the previous moment, and the transient load of the physical machine is large.
[0058] In addition, when the physical machine of the device is operating normally, it may be attacked by external forces. At this time, the attacker may tamper with a certain operation report data reported by the physical machine of the device, causing the transient load value obtained by the current physical machine of the device to be larger. At this time, it is easy to make a misjudgment if the abnormal operation of the physical machine of the device is judged only by the transient load value, which requires further analysis.
[0059] Therefore, when a device physical machine is tampered with by an attacker and exhibits abnormal operation, this is often manifested as a false abnormal increase in a certain operating status data of the device physical machine, resulting in a large difference in the distribution of various operating status data during the operation of the device physical machine compared to the previous moment. At the same time, in order to ensure the stable operation of the service, the same service is often deployed on the same cabinet, and the data center has a load balancing control strategy. Therefore, for a single device physical machine, the smaller the U-position ratio of the cabinet in which it is located, the fewer device physical machines share the load for the service, and the greater the possibility of a real failure. Conversely, if the U-position ratio is larger, it means that there is less empty space on the cabinet, and more device physical machines can be used for load balancing, reducing the operating load of a single device physical machine.
[0060] Based on the above analysis, the abnormal evaluation value of each physical machine of each device at each moment is calculated as follows:
[0061] ;
[0062] Where, Indicates the abnormal evaluation value of the physical machine of the i-th device at the current moment, Indicates the instantaneous load of the physical machine of the i-th device at the current moment, Indicates the range of the deviation value of all running status data of the physical machine of the i-th device at the current moment, Indicates the U-slot ratio of the cabinet where the physical machine of the i-th device is located.
[0063] For devices with fewer physical machines capable of load sharing during operation, and with larger, more uniformly distributed increases in various operational status data, this indicates an overloaded state, increasing the likelihood of an anomaly. Furthermore, for false load increases caused by cyberattacks, the uneven distribution of deviations in various operational status data results in smaller anomaly assessment values, reducing false positives caused by attacks in the data center.
[0064] Step S3: Visually monitor the physical machine of the equipment based on the abnormal evaluation value, and perform alarm maintenance on the faulty physical machine of the equipment.
[0065] The above steps are used to evaluate the operating conditions of all equipment physical machines in the data center at each moment, and obtain the abnormal evaluation value of each equipment physical machine at each moment. The abnormal evaluation values of all equipment physical machines at any moment and all previous moments are used as the input of the cross-validation method, and the output optimal threshold is used as the screening threshold at any moment. The equipment physical machine whose abnormal evaluation value at any moment is greater than the screening threshold is marked as a potential abnormal device at any moment. Only when the equipment physical machine is a potential abnormal device for N consecutive moments is the equipment physical machine marked as a faulty equipment physical machine. It should be noted that the general optional interval for the corresponding time length of the N consecutive moments is [20,60]s. The larger the value, the more accurate the judgment, but the longer the time lag for the processing of the abnormal equipment physical machine. Preferably, in the embodiment of the present application, the value of N is set to 20, that is, when the equipment physical machine is judged as a potential abnormal device for 20 consecutive seconds, it is determined that the equipment physical machine has an operating failure and is marked as a faulty equipment physical machine. Among them, the cross-validation method is a well-known technology, and the specific process will not be repeated.
[0066] Obtain the number of the physical machine of the faulty device and its location in the data center network topology, thereby locating the physical machine of the faulty device in real space, and alert the relevant operation and maintenance personnel to notify them to repair the physical machine of the faulty device.
[0067] Furthermore, in order to facilitate the monitoring of the operation of the entire data center, this application reserves a visualization interface on the intelligent operation and maintenance platform, which dynamically displays the network structure of each device physical machine in the data center, as well as the various operating status data of each device physical machine in real time. If each device physical machine in the data center is marked as a potential abnormal device, the corresponding device physical machine in the visualization interface will be marked yellow; if each device physical machine is marked as a faulty device physical machine, the corresponding device physical machine in the visualization interface will be marked red, and a corresponding operation log will be generated for storage and inspection.
[0068] The process of obtaining the U-space ratio of each cabinet is shown in the following figure: Figure 2 shown.
[0069] In summary, the embodiments of the present application can achieve the expansion of the distributed data center network structure through the network transmission protocol, thereby improving the scalability of data center operation and maintenance; at the same time, compared with the traditional operation and maintenance monitoring process in which the abnormal operation of the equipment physical machine is simply judged by sensors, the present application constructs an abnormal evaluation value for each equipment physical machine based on the correlation between the operation status data of each equipment physical machine in the data center and the U-bit ratio of the equipment physical machine, and evaluates the authenticity of the abnormality of the equipment physical machine; determines the faulty equipment physical machine based on the abnormal evaluation value, and realizes the accurate positioning and timely operation of the abnormal equipment physical machine; avoids the problem of directly judging the operation status of the equipment physical machine by only collecting the operation status data of the equipment physical machine in the data center in the traditional monitoring process, which is easily affected by data transmission interference and false information of network attacks, resulting in misjudgment of the operation status of the equipment physical machine, improves the operation and maintenance efficiency of the traditional data center operation and maintenance method, reduces the risk of misjudgment and missed judgment, and improves the stability and reliability of the data center operation.
[0070] It should be noted that the order in which the embodiments of the present application are presented is for illustrative purposes only and does not necessarily represent the superiority or inferiority of the embodiments. Furthermore, the above descriptions are of specific embodiments of the present application. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential sequence shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0071] The various embodiments in this application are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments.
[0072] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them. Modifications to the technical solutions described in the aforementioned embodiments, or equivalent replacements of some of the technical features therein, do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A data center operation and maintenance visualization monitoring method combining big data and network topology, characterized in that: The method comprises the following steps: For all physical machines in all cabinets in the data center, obtain the operating status of each physical machine at each moment, including temperature, memory usage, CPU usage, and network traffic, normalizing the temperature and network traffic; and obtain the U-slot ratio of each cabinet. The absolute value of the difference between each item of operating status data of each device physical machine at each moment and the previous moment is used as the deviation value of each item of operating status of each device physical machine at each moment; based on the overall distribution characteristics of the deviation values of all items of operating status of each device physical machine at each moment, the transient load of each device physical machine at each moment is calculated, and the transient load is the sum of the deviation values of all items of operating status data of any device physical machine; based on the distribution range of the deviation values of all items of operating status of each device physical machine at each moment, combined with the U-bit proportion and the transient load, the abnormal assessment value of each device physical machine at each moment is calculated; Perform visual monitoring of physical equipment based on abnormal assessment values, and issue alarms and repairs on faulty physical equipment. The process of obtaining the abnormal evaluation value is as follows: Calculate the range of the deviation values of all items of operating status data of each device physical machine at the current moment, and determine the abnormal assessment value of each device physical machine at each moment based on the range of the deviation values, the U-bit proportion and the transient load.
2. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 1, characterized in that: The process of obtaining the U-space ratio of each cabinet is as follows: The U-bit capacity occupied by all physical devices in the cabinet is obtained by the number of physical devices on the cabinet; the U-bit capacity of the cabinet is determined by the container size of the cabinet; and the ratio of the U-bit capacity occupied by all physical devices to the U-bit capacity of the cabinet is used as the U-bit ratio of the cabinet.
3. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 1 is characterized in that: The expression of the abnormal evaluation value of each physical machine of the device at each time is: , where Indicates the abnormal evaluation value of the physical machine of the i-th device at the current moment, Indicates the instantaneous load of the physical machine of the i-th device at the current moment, Indicates the range of the deviation value of the physical machine of the i-th device at the current moment, Indicates the U-slot ratio of the cabinet where the physical machine of the i-th device is located.
4. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 1, characterized in that: The visual monitoring of the physical machine of the device based on the abnormal evaluation value is specifically as follows: Based on the abnormality assessment value, the potential abnormal device is obtained, and based on the continuous occurrence of the potential abnormal device over time, the physical machine of the faulty device is obtained; the potential abnormal device and the physical machine of the faulty device are visually marked.
5. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 4 is characterized in that: The process of obtaining the potential abnormal device is as follows: The abnormal evaluation values of all device physical machines at any moment and all previous moments are used as the input of the cross-validation method, the output optimal threshold is used as the screening threshold at any moment, and the device physical machine whose abnormal evaluation value at any moment is greater than the screening threshold is marked as a potential abnormal device at any moment.
6. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 4, characterized in that: The process of obtaining the physical machine of the faulty device is as follows: If each device physical machine is marked as a potential abnormal device at a preset number of consecutive moments, each device physical machine is marked as a faulty device physical machine.
7. The data center operation and maintenance visualization monitoring method combining big data and network topology according to claim 4 is characterized in that: The visual marking of potential abnormal devices and faulty physical machines is specifically as follows: The intelligent operation and maintenance monitoring platform is connected to the data center through a network line. A visualization interface is reserved on the intelligent operation and maintenance platform. If the physical machines of each device in the data center are marked as potentially abnormal devices, the corresponding physical machines of the device in the visualization interface will be marked yellow. If the physical machines of each device are marked as faulty physical machines, the corresponding physical machines of the device in the visualization interface will be marked red.
Citation Information
Patent Citations
Intelligent CMDB management and cloud host monitoring method in cloud environment
CN109274557A
Intelligent operation and maintenance visual monitoring method and system based on big data
CN117724928A