A Dynamic Monitoring Method, Device, Medium and Program Product for a Server Computer Room

By setting up multi-level alarm policies in the server computer room and dynamically adjusting them in combination with real-time monitoring information, the problem of low monitoring accuracy is solved, efficient, accurate monitoring and rapid response to the server computer room is achieved, and the stable operation and equipment safety of the computer room are ensured.

CN119892680BActive Publication Date: 2025-08-05SHANDONG DENENG IOT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510363185.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-05
Estimated Expiration
2045-03-26

AI Technical Summary

Technical Problem

In the prior art, the monitoring accuracy of the server room is low, making it difficult to effectively capture and respond to dynamically changing high loads and complex environmental interference, resulting in incomplete and inaccurate operating status evaluation.

Method used

By obtaining the historical monitoring information and cluster performance parameters of the server room, setting up multi-level alarm policies, and dynamically adjusting them in combination with internal and external real-time monitoring information, formulating target alarm policies, and performing corresponding alarm processing operations.

Benefits of technology

It improves the monitoring accuracy and response speed of the server computer room, ensures the stability and security of the computer room, and can respond to various abnormal situations in a timely manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119892680B_ABST
    Figure CN119892680B_ABST
Patent Text Reader

Abstract

The present application provides a dynamic monitoring method, device, medium and program product for a server computer room, relating to the technical field of server computer room monitoring. The method includes: when historical monitoring information of the server computer room and cluster performance parameters of a server device cluster are obtained, setting a multi-level alarm policy according to the historical monitoring information and the cluster performance parameters; obtaining internal monitoring information obtained by a first monitoring device for real-time monitoring of the internal environment of the server computer room, and obtaining external monitoring information obtained by a second monitoring device for real-time monitoring of the external environment of the server computer room; performing a dynamic adjustment operation on the multi-level alarm policy according to the internal monitoring information and the external monitoring information to obtain a target alarm policy; and performing an alarm processing operation on the server computer room according to the target alarm policy. The technical problem of low monitoring accuracy in the related art for the server computer room is solved, and the technical effect of improving the monitoring accuracy of the server computer room is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of server room monitoring, and in particular, to a dynamic monitoring method, device, medium and program product for a server room. Background Art

[0002] With the rapid development of cloud computing and big data technologies, as the core infrastructure supporting modern digital services, the stability and reliability of the operating state of server rooms have become key factors in ensuring service quality and business continuity.

[0003] In related technologies, a monitoring and alarm mechanism based on fixed thresholds is generally adopted in server rooms. Specifically, sensors are deployed inside the server room to obtain monitoring data, and this monitoring data is compared with pre-set fixed thresholds. When the monitoring data exceeds the fixed threshold, an alarm signal will be triggered, and the operation and maintenance personnel will be notified through a pre-set communication channel for manual intervention. However, the current monitoring of server rooms needs to cope with multiple challenges such as high-dynamic load changes, complex environmental interference, and device heterogeneity. These challenges make the operating state of server rooms more complex and changeable, and it may be difficult for the monitoring and alarm mechanism based on fixed thresholds to effectively capture and respond to these dynamic changes.

[0004] However, when adopting the above-mentioned monitoring and alarm mechanism based on fixed thresholds, since the workload and environmental conditions of the server room often change, it is difficult for the fixed thresholds to adapt to these dynamic changes, thus unable to comprehensively and accurately evaluate and warn the operating state of the computer room, resulting in a relatively low monitoring accuracy rate in the related technologies for server rooms. Summary of the Invention

[0005] This application provides a dynamic monitoring method, device, medium and program product for a server room, which is used to improve the monitoring accuracy rate of the server room.

[0006] In a first aspect, this application provides a dynamic monitoring method for a server room, which is applied to the above-mentioned electronic device. The method includes: when obtaining the historical monitoring information of the server room and the cluster performance parameters of the server device cluster, setting a multi-level alarm policy according to the historical monitoring information and the cluster performance parameters, where the server room includes a server device cluster; obtaining the internal monitoring information obtained by a first monitoring device for real-time monitoring of the internal environment of the server room, and obtaining the external monitoring information obtained by a second monitoring device for real-time monitoring of the external environment of the server room; performing a dynamic adjustment operation on the multi-level alarm policy according to the internal monitoring information and the external monitoring information to obtain a target alarm policy; and performing an alarm processing operation on the server room according to the target alarm policy.

[0007] By adopting the above technical solution, a multi-level alarm strategy can be set according to historical data and cluster performance parameters, providing a scientific early warning mechanism for the monitoring of the server room. By real-time monitoring of the internal and external environments of the server room, various factors that may affect the stable operation of the server room can be comprehensively captured. The multi-level alarm strategy is dynamically adjusted according to these real-time monitoring information, thus ensuring the timeliness and accuracy of the multi-level alarm strategy. The alarm processing operation is executed according to the adjusted target alarm strategy, which can effectively improve the monitoring efficiency and response speed of the server room. Furthermore, the technical problem of low monitoring accuracy in the server room in the related technology is solved, achieving the technical effect of improving the monitoring accuracy of the server room.

[0008] Optionally, when the historical monitoring information of the server room and the device performance information of the server device cluster are obtained, setting the multi-level alarm strategy according to the historical monitoring information and the cluster performance parameters specifically includes: determining the environmental parameter reference fluctuation range and historical fault time series data of the server room according to the historical monitoring information; determining the temperature, humidity, and pressure tolerance parameter curve and redundancy configuration status of the server device cluster according to the cluster performance parameters; setting the multi-level alarm strategy according to the environmental parameter reference fluctuation range, historical fault time series data, temperature, humidity, and pressure tolerance parameter curve, and redundancy configuration status.

[0009] By adopting the above technical solution, determining the environmental parameter reference fluctuation range and historical fault time series data according to the historical monitoring information provides a data basis for formulating the multi-level alarm strategy. Combining the temperature, humidity, and pressure tolerance parameter curve and redundancy configuration status of the server device cluster can more accurately evaluate the tolerance ability of all devices in the server room and the availability of standby resources. The multi-level alarm strategy set by comprehensively considering various factors not only can improve the pertinence and effectiveness of the alarm strategy, but also can ensure the stable operation of the server room.

[0010] Optionally, set multi-level alarm strategies according to the environmental parameter baseline fluctuation range, historical fault time series data, temperature, humidity, and pressure resistance parameter curves, and redundant configuration status. Specifically, it includes: setting a warning-level strategy according to the environmental parameter baseline fluctuation range and temperature, humidity, and pressure resistance parameter curves. Among them, the warning-level strategy is that when at least one of the three parameters of temperature parameter, humidity parameter, and voltage parameter reaches the first preset percentage range of the device tolerance threshold and deviates from the environmental parameter baseline fluctuation range, trigger fan speed regulation or enable the standby power module; determine the performance degradation data of the server device cluster according to the historical fault time series data, and determine the device load rate of the server device cluster according to the cluster performance parameters; set an alarm-level strategy according to the performance degradation data and device load rate. Among them, the alarm-level strategy is that when at least one of the three parameters of temperature parameter, humidity parameter, and voltage parameter exceeds the device tolerance threshold and the device load rate is greater than the first preset load threshold, trigger the operation of switching to the standby server device or shutting down the non-core business; determine the power supply stability data of the server device cluster according to the cluster performance parameters; set an emergency-level strategy according to the historical fault time series data and power supply stability data. Among them, the emergency-level strategy is that when the three parameters of temperature parameter, humidity parameter, and voltage parameter exceed the preset deviation range of the historical device tolerance peak and the power supply fluctuation amplitude exceeds the safe fluctuation range, trigger the operation of device power-off protection or data emergency migration.

[0011] By adopting the above technical solutions, setting a warning-level strategy according to the environmental parameter baseline fluctuation range and temperature, humidity, and pressure resistance parameter curves can take timely measures when the environmental parameters deviate from the normal range to prevent equipment damage. Setting an alarm-level strategy according to the historical fault time series data and device load rate can timely switch to standby equipment or shut down non-core business when the device performance degrades or the load is too high, thus ensuring the overall performance of the server room. Setting an emergency-level strategy according to the power supply stability data and historical fault time series data can quickly take measures such as power-off protection or data migration in case of unstable power supply or device tolerance limit to ensure data security and device integrity.

[0012] Optionally, obtain the internal monitoring information obtained by the first monitoring device through real-time monitoring of the internal environment of the server room, and obtain the external monitoring information obtained by the second monitoring device through real-time monitoring of the external environment of the server room. Specifically, it includes: obtaining the CPU utilization rate, memory occupancy rate, and network throughput data of the server device cluster collected in real time by the first monitoring device from the internal environment, where the internal monitoring information includes the CPU utilization rate, memory occupancy rate, and network throughput data; obtaining the instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value collected in real time by the first monitoring device from the core area of the internal environment, where the internal monitoring information includes the instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value; obtaining the temperature change trend, air humidity index, and lightning activity warning level collected in real time by the second monitoring device from the external environment, where the external monitoring information includes the temperature change trend, air humidity index, and lightning activity warning level; obtaining the voltage stability coefficient and frequency offset amplitude collected in real time by the second monitoring device from the external power grid in the external environment, where the external monitoring information includes the voltage stability coefficient and frequency offset amplitude, and the voltage stability coefficient is an evaluation index for the power supply quality of the external power grid, and the frequency offset amplitude is the absolute value of the deviation between the actual operating frequency and the standard operating frequency of the external power grid.

[0013] By adopting the above technical solution, various key parameters inside and outside the server room are comprehensively collected, providing rich data support for dynamically adjusting the alarm strategy. The internal monitoring information reflects the operating status and load conditions of the server device cluster. The internal core area monitoring information is directly related to the stability and security of the server room environment. The external environment monitoring information helps to predict potential natural disaster risks in advance. The external power grid monitoring information ensures the reliability and stability of the power supply to the computer room. Through the comprehensive collection and analysis of these information, a comprehensive and accurate data foundation can be provided for the operation and maintenance of the server room.

[0014] Optionally, perform a dynamic adjustment operation on the multi-level alarm strategy according to the internal monitoring information and the external monitoring information to obtain the target alarm strategy. Specifically, it includes: generating a workload intensity evaluation index according to the CPU utilization rate, memory occupancy rate, and network throughput data, and generating an environmental pressure parameter set according to the instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value; generating power supply quality evaluation data according to the voltage stability coefficient and frequency offset amplitude; performing a dynamic correlation analysis on the workload intensity evaluation index, the environmental pressure parameter set, and the power supply quality evaluation data to generate a dynamic monitoring baseline; performing a dynamic adjustment operation on the multi-level alarm strategy according to the dynamic monitoring baseline to obtain the target alarm strategy.

[0015] By adopting the above technical solution, the workload intensity evaluation index, the environmental stress parameter set, and the power supply quality evaluation data are dynamically correlated and analyzed to generate a dynamic monitoring baseline that can reflect the overall operating conditions of the computer room in real time, providing a scientific basis for the dynamic adjustment of the multi-level alarm strategy, thereby ensuring the timeliness and accuracy of the alarm strategy, and further enhancing the security and stability of the server computer room.

[0016] Optionally, perform a dynamic adjustment operation on the multi-level alarm strategy according to the dynamic monitoring baseline to obtain a target alarm strategy, which specifically includes: when it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is lower than the preset voltage stability threshold, reducing the first preset proportion range of the voltage parameter to the second preset proportion range, and increasing the first startup priority of the standby voltage module to the second startup priority, where the second startup priority is higher than the first startup priority; when it is determined according to the dynamic monitoring baseline that the first correlation degree between the workload intensity index and the environmental stress parameter set is greater than the first preset correlation degree threshold, reducing the first preset load threshold to the second preset load threshold; when it is determined according to the dynamic monitoring baseline that the lightning activity warning level is greater than or equal to the preset level threshold and the frequency offset amplitude is greater than the preset amplitude threshold, adjusting the first trigger time of the data emergency migration operation to the second trigger time, where the first trigger time is later than the second trigger time; or, when it is determined according to the dynamic monitoring baseline that the second correlation degree between the workload intensity index and the environmental stress parameter set is less than the second preset correlation degree threshold, adjusting the third trigger time of the data emergency migration operation to the fourth trigger time, where the third trigger time is earlier than the fourth trigger time; when it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is higher than the preset voltage stability threshold, resetting the target alarm strategy.

[0017] By adopting the above technical solution, the alarm strategy is refined and dynamically adjusted according to the dynamic monitoring baseline. When the voltage stability coefficient is lower than the preset voltage stability threshold, the first preset proportion range of the voltage parameter is reduced, and the startup priority of the standby voltage module is increased. The preset load threshold is adjusted according to the correlation degree between the workload intensity index and the environmental stress parameter set, and the trigger time of the data emergency migration operation is adjusted according to the lightning activity warning level and the frequency offset amplitude. In addition, when the voltage stability coefficient is higher than the preset voltage stability threshold, the target alarm strategy is reset. These adjustment measures make the alarm strategy more flexible and effectively enhance the response ability and security of the server computer room.

[0018] Optionally, perform an alarm handling operation on the server room according to the target alarm policy, specifically including: when it is determined that the target alarm policy is a pre-warning level policy, send a voltage fluctuation warning signal and a countdown prompt message for starting the standby power supply; dynamically adjust the air-conditioning cooling power gradient and fan speed level of the server room according to the second preset ratio range; start the standby voltage module according to the second startup priority to complete the seamless switching of the power supply line; or, when it is determined that the target alarm policy is an alarm level policy, continuously monitor the load balancing state of the server device cluster according to the second preset load threshold; when it is determined that the single-node device load rate exceeds the second preset load threshold, trigger the non-core service shutdown operation and generate a resource dynamic allocation policy; when it is determined that the cluster device load rate continuously exceeds the second preset load threshold within a preset time period, trigger the standby server device switching operation and generate a device performance degradation analysis report; store the resource dynamic allocation policy and the device performance degradation analysis report in the server room database; when it is determined that the target alarm policy is an emergency level policy, the first correlation degree is greater than the first preset correlation degree threshold, the lightning activity warning level is greater than or equal to the preset level threshold, and the frequency offset amplitude is greater than the preset amplitude threshold, complete the data emergency migration operation within the second trigger time; after it is determined that the data emergency migration operation is completed, perform the following device power-off protection operations on the server room: cut off the power supply of non-critical devices and release the redundant capacitors of the uninterruptible power supply, where the server room includes non-critical devices; enable surge physical isolation protection for the storage array of the server room; trigger the emergency trip protection of the machine room-level circuit breaker; or, when it is determined that the target alarm policy is an emergency level policy and the second correlation degree is less than the second preset correlation degree threshold, complete the data emergency migration operation within the fourth trigger time.

[0019] By adopting the above technical solutions, when the pre-warning level policy is triggered, by sending a voltage fluctuation warning signal and a countdown prompt message for starting the standby power supply, sufficient preparation time can be provided for relevant operation and maintenance personnel; by dynamically adjusting the air-conditioning cooling power gradient and fan speed level, the stability of the server room environment can be ensured. When the alarm level policy is triggered, by continuously monitoring the load balancing state and triggering corresponding operations, the stable operation of the core business can be ensured. When the emergency level policy is triggered, by quickly completing the data emergency migration and performing the device power-off protection operation, the safety of devices and data can be maximally protected. The comprehensive application of these alarm handling operations can not only ensure that the machine room can quickly and effectively respond to various abnormal situations, but also improve the overall stability and safety of the server room.

[0020] Second aspect, embodiments of the present application provide an electronic device, which includes: one or more processors and a memory; the memory is coupled to the one or more processors, and the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to cause the electronic device to execute the methods described in the first aspect and any possible implementation manner in the first aspect.

[0021] Third aspect, embodiments of the present application provide a computer program product containing instructions, when the above computer program product runs on an electronic device, causing the above electronic device to execute the methods described in the first aspect and any possible implementation manner in the first aspect.

[0022] Fourth aspect, embodiments of the present application provide a computer-readable storage medium, including instructions, when the above instructions run on an electronic device, causing the above electronic device to execute the methods described in the first aspect and any possible implementation manner in the first aspect.

[0023] One or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:

[0024] 1. The dynamic monitoring method of the server room provided by the present application sets a multi-level alarm strategy according to historical data and cluster performance parameters, which can provide a scientific early warning mechanism for the monitoring of the server room. By real-time monitoring of the internal and external environments of the server room, various factors that may affect the stable operation of the server room can be comprehensively captured. Dynamically adjust the multi-level alarm strategy according to these real-time monitoring information, so as to ensure the timeliness and accuracy of the multi-level alarm strategy. Execute the alarm processing operation according to the adjusted target alarm strategy, which can effectively improve the monitoring efficiency and response speed of the server room.

[0025] 2. The dynamic monitoring method of the server room provided by the present application determines the environmental parameter reference fluctuation range and historical fault time series data according to historical monitoring information, providing a data basis for the formulation of the multi-level alarm strategy. Combining the temperature, humidity and pressure tolerance parameter curves and redundant configuration status of the server device cluster, the tolerance capabilities of all devices in the server room and the availability of backup resources can be evaluated more accurately. The multi-level alarm strategy set by comprehensively considering various factors not only can improve the pertinence and effectiveness of the alarm strategy, but also can ensure the stable operation of the server room.

[0026] 3. The dynamic monitoring method for the server room provided by this application sets the early warning level strategy according to the environmental parameter baseline fluctuation range and the temperature, humidity, and pressure resistance parameter curve, and can take timely measures when the environmental parameters deviate from the normal range to prevent equipment damage. The alarm level strategy is set according to the historical failure time series data and the equipment load rate, and can switch to standby equipment or shut down non-core services in a timely manner when the equipment performance degrades or the load is too high, so as to ensure the overall performance of the server room. The emergency level strategy is set according to the power supply stability data and the historical failure time series data, and can quickly take measures such as power-off protection or data migration in case of unstable power supply or the equipment reaching its limit, to ensure the security of data and the integrity of equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 is a schematic flowchart of a dynamic monitoring method for the server room in an embodiment of this application;

[0028] Figure 2 is a schematic structural diagram of an entity device of an electronic device in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] The terms used in the following embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and appended claims of this application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations including one or more of the listed items.

[0030] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as implying or indicating relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise specified, the meaning of "plural" is two or more.

[0031] This application provides a dynamic monitoring method for a server room. Refer to Figure 1 , Figure 1 is a schematic flowchart of a dynamic monitoring method for the server room in an embodiment of this application, including the following steps:

[0032] Step S101, when the historical monitoring information of the server room and the cluster performance parameters of the server device cluster are obtained, set a multi-level alarm strategy according to the historical monitoring information and the cluster performance parameters;

[0033] In the above embodiments, the historical monitoring information represents various types of data collected during a past period (for example, within three days, within a week, within a month, etc., which is not limited here) for monitoring the server room, such as temperature, humidity, power usage, equipment failure records, etc., which is not limited here. The server device cluster refers to a set composed of multiple server devices. The cluster performance parameters refer to a series of indicators describing the overall performance of the server device cluster, such as CPU (Central Processing Unit) usage rate, memory occupancy rate, network throughput, storage I / O (Input / Output) performance, etc., which is not limited here. The multi-level alarm policy refers to different levels of alarm rules set according to the historical monitoring information and the cluster performance parameters.

[0034] In the above embodiments, the execution timing and scenario of step S101 are usually during the operation and maintenance management of the server room to ensure the stable operation of the server device cluster and timely response to potential problems. Specifically, after obtaining the historical monitoring information of the server room and the cluster performance parameters of the server device cluster, the normal operation range and possible abnormal situations of the server room and the server device cluster are analyzed based on the historical monitoring information and the cluster performance parameters. According to these analysis results, a multi-level alarm policy is set, and the multi-level alarm policy includes but is not limited to different levels of alarm thresholds, alarm methods (such as text messages, emails, audible and visual alarms, mobile APP alarms, uploading alarms to the Internet of Things platform, etc.), alarm objects (such as operation and maintenance personnel, management personnel, etc.), and the processing process after the alarm. By setting the multi-level alarm policy, it can be ensured that corresponding processing measures can be taken in a timely and accurate manner when different levels of abnormalities occur.

[0035] Step S102, obtain the internal monitoring information obtained by the first monitoring device for real-time monitoring of the internal environment of the server room, and obtain the external monitoring information obtained by the second monitoring device for real-time monitoring of the external environment of the server room;

[0036] In the above embodiments, the first monitoring device is installed inside the server room and is used to monitor the internal environment of the server room in real time. The first monitoring device can monitor environmental parameters such as temperature, humidity, air quality (e.g., combustible gas, PM2.5, etc.), smoke, water immersion, noise, light collection, digital variable frequency infrared grating, etc. The internal monitoring information refers to various types of data of the internal environment of the server room monitored by the first monitoring device in real time, and these data can reflect the real-time situation inside the server room. The second monitoring device is installed outside the server room and is used to monitor the external environment of the server room in real time. The second monitoring device may monitor natural parameters such as external temperature, humidity, wind speed, rainfall, etc., as well as possible security situations such as intrusion and damage. The external monitoring information refers to various types of data of the external environment of the server room monitored by the second monitoring device in real time, and these data help to understand the environmental changes outside the server room and respond in a timely manner to external factors that may affect the server room.

[0037] In the above embodiments, the execution timing and scenario of step S102 are usually during the daily operation and maintenance of the server room to ensure the stability and safety of the internal and external environments of the server room. Specifically, the first monitoring device and the second monitoring device are configured to monitor the internal and external environments of the server room in real time respectively. The first monitoring device will collect data such as temperature, humidity, and air quality inside the server room regularly or in real time, and save these data as internal monitoring information. At the same time, the second monitoring device will also monitor the environmental parameters and security situations outside the server room and save the collected data as external monitoring information. Through these real-time monitoring information, the real-time situations inside and outside the server room can be understood in a timely manner, potential problems can be discovered, and corresponding measures can be taken for prevention or handling.

[0038] Step S103, perform a dynamic adjustment operation on the multi-level alarm strategy according to the internal monitoring information and the external monitoring information to obtain the target alarm strategy;

[0039] In the above embodiments, the dynamic adjustment operation means that according to the internally monitored information and the externally monitored information obtained in real time, the alarm level, alarm conditions, emergency measures, etc. in the multi-level alarm strategy are adjusted in real time to adapt to the changes in the internal and external environments of the server room. The target alarm strategy refers to the multi-level alarm strategy that is most suitable for the current internal and external environmental conditions of the server room after the dynamic adjustment operation.

[0040] In the above embodiments, the execution timing and scenario of step S103 are usually during the daily operation and maintenance of the server room. When the internal and external environments of the server room change, it is necessary to adjust the alarm policy in real time to ensure the safe and stable operation of the server room. Specifically, according to the internally monitored information and externally monitored information obtained in real time, the environmental conditions inside and outside the server room are analyzed. For example, when it is found that the internal temperature of the server room is gradually rising and the external environmental temperature is also high, it may be determined that the heat dissipation of the server room faces great pressure. At this time, according to the multi-level alarm policy, the alarm level and alarm conditions can be dynamically adjusted, such as lowering the threshold of temperature alarm or increasing the emergency measures after alarm, to obtain the target alarm policy. In this way, when the temperature of the server room rises slightly, the alarm can be triggered in time, and corresponding measures can be taken in time to prevent safety problems such as overheating in the server room.

[0041] Step S104, perform an alarm processing operation on the server room according to the target alarm policy.

[0042] In the above embodiments, the alarm processing operation refers to a series of emergency response measures executed when the monitoring data of the server room triggers the alarm conditions in the target alarm policy. The emergency response measures may include sending alarm information to the operation and maintenance personnel, starting standby equipment, adjusting environmental control parameters, etc., to quickly restore the normal operation state of the server room or prevent the problem from deteriorating further.

[0043] In the above embodiments, the execution timing and scenario of step S104 are usually during the daily operation and maintenance of the server room when the monitoring data of the server room triggers the alarm conditions in the target alarm policy. Specifically, when the monitoring devices (corresponding to the above first monitoring device and the above second monitoring device) real-time monitor the changes in the internal and external environmental parameters of the server room (such as temperature, humidity, power supply, etc., which are not limited here), and these changes trigger the alarm conditions in the target alarm policy, the alarm processing operation will be automatically executed. The alarm processing operation includes but is not limited to immediately notifying the operation and maintenance personnel of the abnormal situation in the server room by means of text message, email or sound, automatically starting the standby power supply, adjusting the air conditioner temperature or opening the smoke exhaust system according to the emergency measures in the alarm policy to quickly restore the normal environment of the server room, recording the alarm information and the processing process to provide a basis for subsequent analysis and improvement, etc.

[0044] Through the above steps, by setting a multi-level alarm policy based on historical data and cluster performance parameters, a scientific early warning mechanism can be provided for the monitoring of the server room. By monitoring the internal and external environments of the server room in real time, various factors that may affect the stable operation of the server room can be comprehensively captured. The multi-level alarm policy is dynamically adjusted according to these real-time monitoring information, so as to ensure the timeliness and accuracy of the multi-level alarm policy. Executing the alarm processing operation according to the adjusted target alarm policy can effectively improve the monitoring efficiency and response speed of the server room. Furthermore, the technical problem of low monitoring accuracy of the server room in the related technology is solved, and the technical effect of improving the monitoring accuracy of the server room is achieved.

[0045] Among them, the execution subject of the above steps can be a control system with the ability to dynamically adjust the alarm policy, or a control device with the ability to dynamically adjust the alarm policy, or a controller or processor in the device or system, or a separately existing controller or processor, or it can also be other processing devices or processing units with similar processing functions, etc., but not limited to this.

[0046] In an optional embodiment, when the historical monitoring information of the server room and the device performance information of the server device cluster are obtained, setting a multi-level alarm policy according to the historical monitoring information and cluster performance parameters specifically includes: determining the environmental parameter baseline fluctuation range and historical fault time series data of the server room according to the historical monitoring information; determining the temperature, humidity, and pressure tolerance parameter curve and redundancy configuration status of the server device cluster according to the cluster performance parameters; setting a multi-level alarm policy according to the environmental parameter baseline fluctuation range, historical fault time series data, temperature, humidity, and pressure tolerance parameter curve, and redundancy configuration status.

[0047] In the above embodiments, assume a data center with multiple server rooms, and a large number of server device clusters are deployed in each server room to process various business applications. To ensure the stable operation of the data center, it is necessary to monitor the environmental parameters and device performance of the server rooms in real time and set multi-level alarm strategies based on this information. The specific implementation steps are as follows: Through sensors and monitoring devices deployed in the server rooms, collect environmental parameter data over a period of time (for example, the past week, the past half month, the past month, etc., not limited here), including but not limited to temperature, humidity, voltage, etc. At the same time, device performance data of the server device clusters is also obtained, such as CPU usage rate, memory occupancy rate, disk I / O, etc. Analyze the collected environmental parameter data to calculate the benchmark fluctuation range of each parameter (i.e., the environmental parameter benchmark fluctuation range), which is the parameter fluctuation range under normal circumstances. Analyze historical fault data to identify the fault time series associated with specific environmental parameter fluctuations (i.e., historical fault time series data), that is, which parameter fluctuations are likely to cause faults. Based on the device performance data of the server device clusters, draw the temperature, humidity, and pressure resistance parameter curves, that is, the performance of the devices under different temperature and humidity conditions. Check the redundancy configuration status of the server devices, including but not limited to power redundancy status, network redundancy status, etc., to ensure that standby devices can be quickly switched to in case of partial device failures. Set multi-level alarm strategies based on the environmental parameter benchmark fluctuation range, historical fault time series data, temperature, humidity, and pressure resistance parameter curves, and redundancy configuration status. Through the above specific implementation steps, reasonable multi-level alarm strategies can be set according to the historical monitoring information of the server rooms and the device performance information of the server device clusters to ensure the stable operation of the data center. In practical applications, the alarm strategies can also be adjusted and optimized according to specific situations to improve the accuracy and timeliness of alarms.

[0048] In an optional embodiment, a multi-level alarm policy is set according to the environmental parameter reference fluctuation range, historical failure time series data, temperature, humidity and pressure resistance parameter curve, and redundancy configuration status, which specifically includes: setting a warning-level policy according to the environmental parameter reference fluctuation range and the temperature, humidity and pressure resistance parameter curve, where the warning-level policy is that when at least one of the three parameters of temperature parameter, humidity parameter, and voltage parameter reaches the first preset proportional range of the device tolerance threshold and deviates from the environmental parameter reference fluctuation range, fan speed regulation is triggered or a standby power supply module is enabled; determining the performance degradation data of the server device cluster according to the historical failure time series data, and determining the device load rate of the server device cluster according to the cluster performance parameters; setting an alarm-level policy according to the performance degradation data and the device load rate, where the alarm-level policy is that when at least one of the three parameters of temperature parameter, humidity parameter, and voltage parameter exceeds the device tolerance threshold and the device load rate is greater than the first preset load threshold, a standby server device switching operation or a non-core service shutdown operation is triggered; determining the power supply stability data of the server device cluster according to the cluster performance parameters; setting an emergency-level policy according to the historical failure time series data and the power supply stability data, where the emergency-level policy is that when the three parameters of temperature parameter, humidity parameter, and voltage parameter exceed the preset deviation range of the historical device tolerance peak value and the power supply fluctuation amplitude exceeds the safe fluctuation range, a device power-off protection operation or a data emergency migration operation is triggered.

[0049] In the above embodiment, assume that there are multiple server rooms in the data center of an Internet company, and a server device cluster is deployed in each server room to support various online services of the company. To ensure the stable operation of the data center, a multi-level alarm policy is set according to factors such as environmental parameters, device performance, and power supply stability. The specific implementation steps are as follows: Environmental parameter data such as temperature, humidity, and voltage are collected in real time through temperature and humidity sensors and voltage monitoring devices in the server room. Determine the reference fluctuation range of temperature, humidity, and voltage. For example, the reference fluctuation range is temperature 20 - 25°C, humidity 40% - 60%, and voltage 220 ± 5V. Draw a temperature, humidity and pressure resistance parameter curve to clarify the performance threshold of the server device under different temperature and humidity conditions. For example, the performance of the server device may be affected when the temperature exceeds 32°C or the humidity exceeds 75%. When any one of the temperature, humidity, and voltage reaches the first preset proportional range (for example, 90%, and of course it can also be 93%, 95%, 98%, etc., which is not limited here) of the device tolerance threshold (for example, temperature 35°C, humidity 80%, voltage 230V or 210V, etc.), that is, the temperature reaches 31.5°C, the humidity reaches 72%, the voltage reaches 225V or 215V, and deviates from the reference fluctuation range, the warning-level policy is triggered, that is, the fan speed in the computer room is automatically adjusted to strengthen heat dissipation, or a standby power supply module is enabled to ensure the stability of power supply.

[0050] In the above embodiments, by analyzing historical failure time series data, a pattern is found that the performance of server devices gradually degrades after long-term high-load operation. By collecting in real time the device load rates of server devices, such as CPU usage rate, memory occupancy rate, etc. When any one of the parameters of temperature, humidity, and voltage exceeds the device tolerance threshold (for example, temperature 35°C, humidity 80%, voltage 230V or 210V, etc.), and the device load rate is greater than the first preset load threshold (for example, CPU usage rate exceeds 75%, memory occupancy rate exceeds 65%, etc.), an alarm-level policy is triggered, that is, automatically switch to a standby server device to replace the device with degraded performance or a fault, or shut down some non-core services to reduce the server load.

[0051] In the above embodiments, the power supply stability data of the server device cluster is collected in real time, such as voltage fluctuation amplitude, frequency stability, etc. Combining with historical failure time series data, the correlation between power supply fluctuations and serious device failures is analyzed. When any one of the parameters of temperature, humidity, and voltage exceeds the preset deviation range of the historical device tolerance peak (for example, temperature exceeds 38°C, humidity exceeds 85%, voltage fluctuation exceeds ±10%, etc.), and the power supply fluctuation amplitude exceeds the safe fluctuation range (for example, voltage fluctuation exceeds ±5%, etc.), an emergency-level policy is triggered, that is, immediately disconnect the power supply of the affected server device to protect the server device from damage, and at the same time start an emergency data migration operation to migrate key data to other secure storage devices or data centers. Through the above specific implementation steps, the data center can effectively monitor and manage the environmental parameters, device performance, and power supply stability in the server computer room. When an abnormal situation occurs, it can be detected in time and corresponding measures can be taken to ensure the stable operation of the data center and the safety of the devices. At the same time, the setting of the multi-level alarm policy can also be flexibly adjusted and optimized according to the actual situation to adapt to changes in different scenarios and requirements.

[0052] In an optional embodiment, obtaining internal monitoring information obtained by a first monitoring device through real-time monitoring of the internal environment of a server room, and obtaining external monitoring information obtained by a second monitoring device through real-time monitoring of the external environment of the server room, specifically includes: obtaining the CPU utilization rate, memory occupancy rate, and network throughput data of a server device cluster collected by the first monitoring device in real time from the internal environment, where the internal monitoring information includes the CPU utilization rate, memory occupancy rate, and network throughput data; obtaining the instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value collected by the first monitoring device in real time from the core area of the internal environment, where the internal monitoring information includes the instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value; obtaining the temperature change trend, air humidity index, and lightning activity warning level collected by the second monitoring device in real time from the external environment, where the external monitoring information includes the temperature change trend, air humidity index, and lightning activity warning level; obtaining the voltage stability coefficient and frequency deviation amplitude collected by the second monitoring device in real time from the external power grid of the external environment, where the external monitoring information includes the voltage stability coefficient and frequency deviation amplitude, the voltage stability coefficient is an evaluation index for the power supply quality of the external power grid, and the frequency deviation amplitude is the absolute value of the deviation between the actual operating frequency and the standard operating frequency of the external power grid.

[0053] In the above embodiments, the device type of the first monitoring device can be an integrated environmental monitoring sensor array, including but not limited to CPU utilization sensors, memory occupancy sensors, network throughput sensors, temperature and humidity sensors, voltage sensors, etc. It can be set to collect data on the CPU utilization, memory occupancy, and network throughput of the server device cluster once per minute (of course, it can also be once every 10 seconds, 30 seconds, 50 seconds, etc., which is not limited here), and collect the instantaneous temperature fluctuation value, humidity instantaneous fluctuation value, and power supply voltage instantaneous fluctuation value in the core area of the computer room once every 5 minutes (of course, it can also be once every 1 minute, 3 minutes, 4 minutes, etc., which is not limited here). The device type of the second monitoring device can be a combined device of a weather station and a power grid monitoring station, including but not limited to temperature sensors, humidity sensors, lightning detectors, voltage monitoring modules, frequency monitoring modules, etc. It can be set to update the temperature change trend, air humidity index, and lightning activity warning level of the external environment once every 10 minutes (of course, it can also be once every 3 minutes, 5 minutes, 7 minutes, etc., which is not limited here), and obtain the voltage stability coefficient and frequency deviation amplitude of the external power grid once per hour (of course, it can also be once every 15 minutes, 30 minutes, 40 minutes, etc., which is not limited here). All the collected data is transmitted to the data processing center of the data center through wired or wireless networks for real-time analysis to identify abnormal trends or potential problems. When the CPU utilization or memory occupancy exceeds the preset threshold, the load balancing or resource scheduling strategy is automatically triggered to optimize the server performance. When the instantaneous fluctuation value of temperature or humidity exceeds the normal range, the air conditioner or dehumidification device in the server computer room is automatically adjusted to maintain a suitable operating environment. When the instantaneous fluctuation value of the power supply voltage is too large, the backup power supply or uninterruptible power supply is started to ensure the stability of the power supply. When the temperature change trend indicates that high-temperature weather is about to occur, the cooling system of the server computer room is started in advance to prevent overheating. When the air humidity index is too high, the dehumidification measures in the server computer room are strengthened to prevent damage to the equipment caused by humidity. When a lightning activity warning is received, the lightning protection facilities in the server computer room are immediately inspected and strengthened to ensure the safety of the equipment. When the voltage stability coefficient decreases or the frequency deviation amplitude increases, coordination is carried out with the power supplier to ensure the power supply quality of the external power grid and prepare an emergency power supply plan. The data of all monitoring devices is integrated into a unified monitoring platform to achieve centralized management and visual display of the data. The monitoring platform provides functions such as real-time data charts, historical data trend analysis, and abnormal event alarms, facilitating the operation and maintenance personnel to quickly understand the internal and external environment conditions of the computer room and make corresponding decisions. Through the above specific implementation steps: the data center can comprehensively and real-time understand the operating environment of the computer room, timely discover and address potential problems, and ensure the stable operation of the server device cluster.

[0054] In an optional embodiment, a dynamic adjustment operation is performed on the multi-level alarm policy according to internal monitoring information and external monitoring information to obtain a target alarm policy, which specifically includes: generating a workload intensity evaluation index based on CPU utilization rate, memory occupancy rate, and network throughput data, and generating an environmental pressure parameter set based on instantaneous temperature fluctuation value, instantaneous humidity fluctuation value, and instantaneous power supply voltage fluctuation value; generating power supply quality evaluation data based on voltage stability coefficient and frequency offset amplitude; performing dynamic correlation analysis on the workload intensity evaluation index, the environmental pressure parameter set, and the power supply quality evaluation data to generate a dynamic monitoring baseline; and performing a dynamic adjustment operation on the multi-level alarm policy according to the dynamic monitoring baseline to obtain a target alarm policy.

[0055] In the above embodiment, taking the monitoring of a cabinet in an industrial data center as an example, the steps for generating the dynamic monitoring baseline are as follows: The CPU utilization rate of the embedded industrial control computer in the cabinet is collected once every 5 seconds, and the value is an integer value between 0 and 100%. The memory occupancy rate is recorded as the percentage of the used capacity / total capacity, and is also collected once every 5 seconds. The network throughput is statistically recorded as megabytes per second (MB / s), and the data is summarized every 5 seconds. One temperature and humidity sensor is deployed at the front and rear ends of the cabinet respectively. The instantaneous temperature fluctuation value is recorded with a period of 250 ms, and the absolute value of the temperature difference between adjacent sampling points is taken. The instantaneous humidity fluctuation value is calculated as the relative humidity change rate, and is also recorded with a period of 250 ms. The instantaneous power supply voltage fluctuation value is measured as the peak-to-peak deviation of the AC voltage and is recorded with a period of 250 ms. Power monitoring continuously obtains data through a smart meter. The voltage stability coefficient is calculated as the reciprocal of the voltage standard deviation within a 10-minute window, and the frequency offset amplitude is recorded as the difference in parts per million from the 50 Hz reference value. The workload intensity evaluation index is set such that when the CPU utilization rate > 80% and the memory occupancy rate > 75%, it is in a high-load state. When the network throughput continuously exceeds 10 MB / s for 30 seconds, the load intensity level is triggered to increase, and it is divided into three levels: low, medium, and high. The environmental pressure parameter set is that when the temperature fluctuation exceeds ±2 °C / second continuously for 5 times, or the humidity fluctuation > ±5%RH / second, or the voltage fluctuation > ±10% of the rated value, an environmental anomaly mark is generated. The power quality evaluation data is that when the voltage stability coefficient is lower than 0.95, it is determined to be an unstable power supply state, and when the frequency offset exceeds ±0.5 Hz, it is marked as a power anomaly.

[0056] As shown in Table 1, set the comprehensive weights of load intensity, environmental status, and power quality:

[0057] Load intensity Environmental status Power quality Comprehensive weight Low Normal Excellent 0.3 Medium 1 item of abnormality Minor fluctuation 0.6 High ≥2 items of abnormality Severe deterioration 0.9

[0058] Table 1 Three-dimensional correlation weight table

[0059] According to the three-dimensional correlation weight table, calculate the dynamic score:

[0060] Score = Load Level Coefficient × 40% + Number of Environmental Abnormalities × 30% + Degree of Power Deterioration × 30%

[0061] Among them, the load level coefficient is determined according to the load intensity level (for example, low = 1, medium = 2, high = 3, etc.). The number of environmental abnormalities is determined according to the number of abnormal items in the environmental pressure parameter set. The degree of power deterioration is determined according to the power quality assessment data (for example, excellent = 0, slight fluctuation = 1, severe deterioration = 2, etc.).

[0062] Dynamic Monitoring Baseline Generation:

[0063] When the score in three consecutive calculation cycles is ≥ 0.7, a dynamic monitoring baseline is generated, and the temperature fluctuation threshold is tightened to ±1.5 °C / second, and the load overload determination time window is shortened to 15 seconds.

[0064] The initial multi-level alarm strategy is set as the early warning level strategy, alarm level strategy, emergency level strategy, etc. The alarm thresholds and alarm levels in the multi-level alarm strategy can be adjusted according to the dynamic monitoring baseline. When the dynamic monitoring baseline changes, the multi-level alarm strategy is adjusted according to the new dynamic monitoring baseline. For example, after the temperature fluctuation threshold is tightened, if the temperature fluctuation exceeds the new threshold, a higher-level alarm is triggered. After the load overload determination time window is shortened, if the network throughput increases sharply in a short period of time, the alarm is triggered faster.

[0065] Through the above specific implementation steps, the multi-level alarm strategy can be dynamically adjusted according to the actual operating status and environmental conditions of the cabinets in the industrial data center, so as to more accurately reflect the operating conditions of the cabinets and timely discover potential problems.

[0066] In an optional embodiment, a dynamic adjustment operation is performed on a multi-level alarm policy according to a dynamic monitoring baseline to obtain a target alarm policy, which specifically includes: when it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is lower than a preset voltage stability threshold, narrowing the first preset proportional range of voltage parameters to a second preset proportional range, and raising the first startup priority of the standby voltage module to a second startup priority, where the second startup priority is higher than the first startup priority; when it is determined according to the dynamic monitoring baseline that the first correlation degree between the workload intensity index and the environmental pressure parameter set is greater than a first preset correlation degree threshold, lowering the first preset load threshold to a second preset load threshold; when it is determined according to the dynamic monitoring baseline that the lightning activity warning level is greater than or equal to a preset level threshold and the frequency offset amplitude is greater than a preset amplitude threshold, adjusting the first trigger time of the data emergency migration operation to a second trigger time, where the first trigger time is later than the second trigger time; or, when it is determined according to the dynamic monitoring baseline that the second correlation degree between the workload intensity index and the environmental pressure parameter set is less than a second preset correlation degree threshold, adjusting the third trigger time of the data emergency migration operation to a fourth trigger time, where the third trigger time is earlier than the fourth trigger time; when it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is higher than the preset voltage stability threshold, resetting the target alarm policy.

[0067] In the above embodiment, assume that in the server room of a large data center, in order to ensure the stable operation of the data center and the security of data, the multi-level alarm policy is dynamically adjusted according to the baseline. The specific implementation steps are as follows: Set key monitoring indicators such as voltage stability coefficient, workload intensity index, environmental pressure parameter set, and lightning activity warning level. Set the preset thresholds and correlation degree thresholds for each indicator. Continuously monitor the voltage stability coefficient. When the voltage stability coefficient is lower than the preset voltage stability threshold (for example, 0.95, 0.97, 0.98, etc., which are not limited here). Narrow the first preset proportional range of voltage parameters (for example, ±5%, etc.) to a second preset proportional range (for example, ±1%, ±2%, ±3%, etc., which are not limited here). Raise the first startup priority of the standby voltage module to a second startup priority (for example, from "standby" to "emergency"). Continuously monitor the correlation degree between the workload intensity index and the environmental pressure parameter set. When the first correlation degree is greater than the first preset correlation degree threshold (for example, 0.8, 0.85, 0.9, etc., which are not limited here), lower the first preset load threshold (for example, CPU utilization rate of 80%, etc.) to a second preset load threshold (for example, CPU utilization rate of 70%, etc.).

[0068] In the above embodiments, the lightning activity warning level and the frequency offset amplitude are continuously monitored. When the lightning activity warning level is greater than or equal to a preset level threshold (e.g., severe) and the frequency offset amplitude is greater than a preset amplitude threshold (e.g., ±0.5 Hz), the first trigger time (e.g., 30 minutes) of the data emergency migration operation is adjusted to the second trigger time (e.g., 15 minutes). When the second correlation degree between the workload intensity index and the environmental stress parameter set is less than the second preset correlation degree threshold (e.g., 0.2) (indicating that the relationship between the load and the environmental stress is relatively loose, which may mean there is a relatively large buffer space), the third trigger time (e.g., 10 minutes) of the data emergency migration operation is adjusted to the fourth trigger time (e.g., 20 minutes). The voltage stability coefficient is continuously monitored. When the voltage stability coefficient is higher than the preset voltage stability threshold, the target alarm strategy is reset to the initial state or re-set according to a new baseline. Through the above specific implementation steps, the multi-level alarm strategy is dynamically adjusted according to the dynamic monitoring baseline, realizing the refined management and optimization of the data center. By real-time monitoring of key indicators and adjusting the alarm strategy according to preset conditions, it is possible to more effectively respond to changes and abnormal situations in the server room, improving the stability of the server room and the security of data.

[0069] In an optional embodiment, an alarm processing operation is performed on the server computer room according to the target alarm policy, which specifically includes: when it is determined that the target alarm policy is the early warning level policy, sending a voltage fluctuation early warning signal and a standby power supply start countdown prompt message; dynamically adjusting the air-conditioning refrigeration power gradient and fan speed level of the server computer room according to the second preset ratio range; starting the standby voltage module according to the second start priority to complete the seamless switching of the power supply line; or, when it is determined that the target alarm policy is the alarm level policy, real-time monitoring the load balancing state of the server device cluster according to the second preset load threshold; when it is determined that the single-node device load rate exceeds the second preset load threshold, triggering a non-core service shutdown operation and generating a resource dynamic allocation policy; when it is determined that the cluster device load rate continuously exceeds the second preset load threshold within a preset time period, triggering a standby server device switching operation and generating a device performance degradation analysis report; storing the resource dynamic allocation policy and the device performance degradation analysis report in the server computer room database; when it is determined that the target alarm policy is the emergency level policy, the first correlation degree is greater than the first preset correlation degree threshold, the lightning activity early warning level is greater than or equal to the preset level threshold, and the frequency deviation amplitude is greater than the preset amplitude threshold, completing the data emergency migration operation within the second trigger time; after it is determined that the data emergency migration operation is completed, performing the following device power-off protection operations on the server computer room: cutting off the power supply of non-critical devices and releasing the redundant capacitors of the uninterruptible power supply, where the server computer room includes non-critical devices; enabling surge physical isolation protection for the storage array of the server computer room; triggering the emergency trip protection of the computer room-level circuit breaker; or, when it is determined that the target alarm policy is the emergency level policy and the second correlation degree is less than the second preset correlation degree threshold, completing the data emergency migration operation within the fourth trigger time.

[0070] In the above embodiment, it is assumed that in the server computer room of a large data center, in order to ensure the stable operation of the server computer room and the security of data, corresponding alarm processing operations need to be performed according to the target alarm policy. The specific implementation steps are as follows: It is monitored that the voltage fluctuation is close to the early warning threshold, and it is determined that the target alarm policy is the early warning level policy. A voltage fluctuation early warning signal is sent to the computer room management personnel to prompt the possible upcoming voltage instability situation, and at the same time, a standby power supply start countdown prompt message is sent to inform that the standby power supply will automatically start after the countdown ends. According to the second preset ratio range, the air-conditioning refrigeration power gradient of the server computer room is dynamically adjusted to adapt to possible temperature changes, and the fan speed level is adjusted to ensure the air circulation and heat dissipation effect in the computer room. According to the second start priority, the standby voltage module is started to achieve seamless switching of the power supply line and ensure the continuous power supply of the server devices.

[0071] In the above embodiment, when it is detected that the load balancing state of the server device cluster is abnormal, the target alarm policy is determined to be the alarm level policy. According to the second preset load threshold, the load balancing state of the server device cluster is monitored in real time. After it is determined that the load rate of a single-node device exceeds the second preset load threshold, a non-core service shutdown operation is triggered to release resources, a resource dynamic allocation policy is generated to optimize resource allocation, and the normal operation of key services is ensured. After it is determined that the load rate of the cluster devices continuously exceeds the second preset load threshold within a preset time period, a standby server device switching operation is triggered, a device performance degradation analysis report is generated, the reasons for the device performance decline are analyzed, and a basis is provided for subsequent maintenance and upgrade. The resource dynamic allocation policy and the device performance degradation analysis report are stored in the server room database for subsequent query and analysis.

[0072] In the above embodiment, the emergency level policy processing (Case 1): It is detected that the first correlation degree is greater than the first preset correlation degree threshold, the lightning activity warning level is greater than or equal to the preset level threshold, and the frequency offset amplitude is greater than the preset amplitude threshold. The target alarm policy is determined to be the emergency level policy. Within the second trigger time, the data emergency migration operation is completed, and the data is migrated to a safe and reliable storage location. The device power-off protection is executed, that is, after it is determined that the data emergency migration operation is completed, the power supply of non-critical devices is cut off, the redundant capacitor of the uninterruptible power supply is released, the surge protection of the storage array in the server room is enabled, and the emergency tripping protection of the room-level circuit breaker is triggered to ensure the safety of the server device cluster. The emergency level policy processing (Case 2): It is detected that the second correlation degree is less than the second preset correlation degree threshold, and the target alarm policy is determined to be the emergency level policy. Within the fourth trigger time, the data emergency migration operation is completed to ensure the security of the data. Through the above specific implementation steps, the corresponding alarm processing operations are performed on the server room according to the target alarm policy, realizing the refined management and optimization of the server device cluster. By monitoring key indicators in real time and performing corresponding operations according to preset conditions, it is possible to more effectively respond to changes and anomalies in the server room, improve the stability of the server room and the security of data. At the same time, by generating a resource dynamic allocation policy and a device performance degradation analysis report, strong support is provided for subsequent maintenance and upgrade.

[0073] Through the embodiments of the present application, dynamic adjustment operations are performed on the preset multi-level alarm policy according to the internally monitored information and externally monitored information obtained in real time. This means that the thresholds of the alarm policy are no longer fixed, but can be adjusted in real time according to environmental changes to more accurately reflect the actual operating conditions of the computer room. This dynamic threshold adjustment mechanism can capture subtle environmental changes, whether it is a small internal temperature fluctuation or a short-term external power instability, which can be detected in time, thereby improving the timeliness and accuracy of early warning.

[0074] The electronic device in the embodiment of the present invention application will be described from the perspective of hardware processing. Refer to Figure 2 , Figure 2 which is a schematic structural diagram of an entity device of the electronic device in the embodiment of the present application.

[0075] It should be noted that Figure 2 the structure of the electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present invention.

[0076] As Figure 2 shown, the electronic device includes a central processing unit (CPU) 201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 202 or the program loaded from the storage part 208 into the random access memory (RAM) 203, such as executing the method described in the above embodiments. In the RAM 203, various programs and data required for system operation are also stored

[0077] . The CPU 201, ROM 202, and RAM 203 are connected to each other via a bus 204. The input / output (I / O) interface 205 is also connected to the bus 204.

[0078] The following components are connected to the I / O interface 205: an input part 206 including an audio input device, a button switch, etc.; an output part 207 including a liquid crystal display (LCD), an audio output device, an indicator light, etc.; a storage part 208 including a hard disk, etc.; and a communication part 209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication part 209 performs communication processing via a network such as the Internet. The drive 210 is also connected to the I / O interface 205 as required. A removable medium 211, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 210 as required, so that the computer program read from it can be installed into the storage part 208 as required.

[0079] Specifically, according to an embodiment of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through the communication part 209, and / or installed from the removable medium 211. When the computer program is executed by the central processing unit (CPU) 201, various functions defined in the present invention are executed.

[0080] It should be noted that specific examples of the computer-readable storage medium may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fibers, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.

[0081] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above module, program segment, or part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block may occur in a different order than that marked in the accompanying drawings.

[0082] Specifically, the electronic device of this embodiment includes a processor and a memory, and a computer program is stored on the memory. When the computer program is executed by the processor, the dynamic monitoring method of the server room provided in the above embodiment is implemented.

[0083] On the other hand, the present invention also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist separately and not be assembled into the electronic device. The above storage medium carries one or more computer programs. When the above one or more computer programs are executed by a processor of the electronic device, the electronic device implements the dynamic monitoring method of the server room provided in the above embodiment.

[0084] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than limiting it; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present application.

[0085] Those of ordinary skill in the art can understand all or part of the processes in the methods of the above embodiments. This process can be completed by relevant hardware instructed by a computer program. This program can be stored in a computer-readable storage medium. When this program is executed, it can include the processes of the above method embodiments. The foregoing storage media include: various media such as ROM or random access memory RAM, magnetic disks, or optical discs that can store program codes.

Claims

1. A dynamic monitoring method for a server room, characterized in that: include: Upon obtaining historical monitoring information of a server room and cluster performance parameters of a server device cluster, setting a multi-level alarm strategy according to the historical monitoring information and the cluster performance parameters, wherein the server room includes the server device cluster; Acquiring internal monitoring information obtained by real-time monitoring of the internal environment of the server room by a first monitoring device, and acquiring external monitoring information obtained by real-time monitoring of the external environment of the server room by a second monitoring device; Performing a dynamic adjustment operation on the multi-level alarm strategy according to the internal monitoring information and the external monitoring information to obtain a target alarm strategy; Executing an alarm processing operation on the server room according to the target alarm strategy; The dynamically adjusting the multi-level alarm strategy according to the internal monitoring information and the external monitoring information to obtain a target alarm strategy specifically includes: Generating a workload intensity assessment index based on the CPU utilization, memory occupancy, and network throughput data included in the internal monitoring information, and generating an environmental pressure parameter set based on the instantaneous temperature fluctuation value, the instantaneous humidity fluctuation value, and the instantaneous power supply voltage fluctuation value included in the internal monitoring information; generating power supply quality assessment data based on the voltage stability coefficient and frequency deviation amplitude included in the external monitoring information; Performing dynamic correlation analysis on the workload intensity assessment indicator, the environmental pressure parameter set, and the power supply quality assessment data to generate a dynamic monitoring baseline; Performing the dynamic adjustment operation on the multi-level alarm strategy according to the dynamic monitoring baseline to obtain the target alarm strategy; Dynamically correlating and analyzing the workload intensity assessment indicator, the environmental pressure parameter set, and the power supply quality assessment data to generate a dynamic monitoring baseline, specifically including: Determining a load level coefficient, a number of environmental anomalies, and a degree of power degradation based on the workload intensity assessment index, the environmental pressure parameter set, and the power supply quality assessment data; Dynamic scoring is performed according to the load level coefficient, the number of environmental abnormalities, and the degree of power degradation, so as to generate the dynamic monitoring baseline when it is determined that three consecutive dynamic scores are greater than or equal to a preset score value.

2. The method according to claim 1, characterized in that When the historical monitoring information of the server room and the equipment performance information of the server equipment cluster are obtained, setting a multi-level alarm strategy according to the historical monitoring information and the cluster performance parameters specifically includes: Determining the baseline fluctuation range of the environmental parameters of the server room and historical failure time series data based on the historical monitoring information; Determining a temperature, humidity and pressure resistance parameter curve and a redundant configuration state of the server device cluster according to the cluster performance parameters; The multi-level alarm strategy is set according to the environmental parameter benchmark fluctuation range, the historical fault time series data, the temperature, humidity and pressure resistance parameter curve, and the redundant configuration status.

3. The method according to claim 2, characterized in that The multi-level alarm strategy is set according to the environmental parameter reference fluctuation range, the historical fault time series data, the temperature, humidity and pressure resistance parameter curve, and the redundant configuration status, specifically including: An early warning level strategy is set according to the environmental parameter baseline fluctuation range and the temperature, humidity and pressure resistance parameter curve, wherein the early warning level strategy is to trigger fan speed regulation or enable the backup power module when at least one of the temperature parameter, humidity parameter and voltage parameter reaches a first preset proportion range of the device tolerance threshold and deviates from the environmental parameter baseline fluctuation range; determining performance degradation data of the server device cluster based on the historical failure time series data, and determining a device load rate of the server device cluster based on the cluster performance parameters; Setting an alarm level strategy based on the performance degradation data and the device load rate, wherein the alarm level strategy triggers a standby server device switching operation or a non-core service shutdown operation when at least one of the temperature parameter, the humidity parameter, and the voltage parameter exceeds the device tolerance threshold and the device load rate is greater than a first preset load threshold; Determining power supply stability data of the server device cluster based on the cluster performance parameters; An emergency level strategy is set based on the historical fault time series data and the power supply stability data, wherein the emergency level strategy is to trigger a device power-off protection operation or a data emergency migration operation when the temperature parameter, humidity parameter, and voltage parameter exceed the preset deviation range of the historical equipment's peak value and the power supply fluctuation amplitude exceeds the safe fluctuation range.

4. The method according to claim 1, wherein The obtaining of internal monitoring information obtained by real-time monitoring of the internal environment of the server room by the first monitoring device, and the obtaining of external monitoring information obtained by real-time monitoring of the external environment of the server room by the second monitoring device, specifically include: Obtaining CPU utilization, memory occupancy, and network throughput data of the server device cluster collected in real time by the first monitoring device from the internal environment; Obtaining instantaneous temperature fluctuation values, instantaneous humidity fluctuation values, and instantaneous power supply voltage fluctuation values collected in real time by the first monitoring device from the core area of the internal environment; Obtaining temperature change trends, air humidity index, and lightning activity warning levels collected in real time by the second monitoring device from the external environment, wherein the external monitoring information includes the temperature change trends, the air humidity index, and the lightning activity warning levels; Obtain the voltage stability coefficient and frequency offset amplitude collected in real time by the second monitoring device from the external power grid of the external environment, wherein the voltage stability coefficient is a power supply quality assessment indicator of the external power grid, and the frequency offset amplitude is the absolute value of the deviation between the actual operating frequency of the external power grid and the standard operating frequency.

5. The method according to any one of claims 1 to 4, characterized in that Performing the dynamic adjustment operation on the multi-level alarm strategy according to the dynamic monitoring baseline to obtain the target alarm strategy specifically includes: When it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is lower than a preset voltage stability threshold, narrowing the first preset proportion range of the voltage parameter to a second preset proportion range, and raising the first startup priority of the backup voltage module to a second startup priority, wherein the second startup priority is higher than the first startup priority; When it is determined according to the dynamic monitoring baseline that the first correlation between the workload intensity index and the environmental pressure parameter set is greater than the first preset correlation threshold, the first preset load threshold is lowered to the second preset load threshold; when it is determined according to the dynamic monitoring baseline that the lightning activity warning level is greater than or equal to the preset level threshold and the frequency offset amplitude is greater than the preset amplitude threshold, the first trigger time of the data emergency migration operation is adjusted to the second trigger time, wherein the first trigger time is later than the second trigger time; or When it is determined according to the dynamic monitoring baseline that a second correlation between the workload intensity indicator and the environmental pressure parameter set is less than a second preset correlation threshold, adjusting the third trigger time of the data emergency migration operation to a fourth trigger time, wherein the third trigger time is earlier than the fourth trigger time; When it is determined according to the dynamic monitoring baseline that the voltage stability coefficient is higher than the preset voltage stability threshold, the target alarm strategy is reset.

6. The method according to claim 5, characterized in that The performing of an alarm processing operation on the server room according to the target alarm strategy specifically includes: When it is determined that the target alarm strategy is a warning-level strategy, a voltage fluctuation warning signal and a backup power startup countdown prompt are sent; the air conditioning cooling power gradient and fan speed level of the server room are dynamically adjusted according to the second preset ratio range; the backup voltage module is started according to the second startup priority to complete seamless switching of the power supply line; or If it is determined that the target alarm policy is an alarm-level policy, the load balancing status of the server device cluster is monitored in real time according to the second preset load threshold; if it is determined that the load rate of a single-node device exceeds the second preset load threshold, a non-core business shutdown operation is triggered, and a resource dynamic allocation policy is generated; if it is determined that the cluster device load rate continues to exceed the second preset load threshold for a preset period of time, a standby server device switching operation is triggered, and a device performance degradation analysis report is generated; the resource dynamic allocation policy and the device performance degradation analysis report are stored in a server room database; When it is determined that the target alarm strategy is an emergency-level strategy, the first correlation is greater than the first preset correlation threshold, the lightning activity warning level is greater than or equal to the preset level threshold, and the frequency offset amplitude is greater than the preset amplitude threshold, the data emergency migration operation is completed within the second trigger time; after determining that the data emergency migration operation is completed, the following equipment power-off protection operations are performed on the server room: cutting off the power supply to non-critical equipment and releasing the redundant capacitor of the uninterruptible power supply, wherein the server room includes the non-critical equipment; enabling surge protection of the storage array of the server room; triggering the emergency trip protection of the room-level circuit breaker; or, When it is determined that the target alarm policy is an emergency-level policy and the second correlation is less than the second preset correlation threshold, the data emergency migration operation is completed within the fourth trigger time.

7. An electronic device, characterized in that: The electronic device includes: one or more processors and a memory; the memory is coupled to the one or more processors, the memory is used to store computer program code, the computer program code includes computer instructions, and the one or more processors call the computer instructions to enable the electronic device to execute the method as described in any one of claims 1-6.

8. A computer-readable storage medium comprising instructions, characterized in that: When the instructions are executed on an electronic device, the electronic device is caused to execute the method according to any one of claims 1 to 6.

9. A computer program product, characterized in that When the computer program product is run on an electronic device, the electronic device is enabled to perform the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Real-time monitoring system for cloud computing server room

    CN118509248A

  • Server hardware fault early warning and recovery method and system based on intelligent optimization algorithm

    CN119621442A