Intelligent operation and maintenance monitoring system for machine room

Through the integrated status monitoring, abnormal warning, data acquisition and load evaluation modules, multi-dimensional intelligent management of the computer room server is realized, solving the problems of insufficient data analysis and lagging response in traditional monitoring systems, and ensuring the stable operation and business continuity of the server.

CN120434280APending Publication Date: 2025-08-05LIDERSHIP TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510221938.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The traditional computer room server operation and maintenance monitoring system lacks multi-dimensional data processing capabilities and intelligent analysis, which makes it difficult to sense server abnormalities in time, which can easily cause equipment overload and hardware loss, and it is difficult to achieve effective early warning, affecting business continuity and security.

Method used

The status monitoring module is used to monitor the temperature distribution and vibration status of the server surface in real time, and combine the abnormal warning module to conduct multi-dimensional data correlation analysis. The data acquisition module is used to obtain memory and network transmission status data, and the load evaluation module is used to build a load degree evaluation model, and intelligent control instructions are generated through the control module to achieve accurate judgment and management of the server load status.

Benefits of technology

It realizes all-round perception and intelligent management of computer room servers, improves the sensitivity and accuracy of abnormal detection, shortens the problem monitoring time, ensures the stability and business continuity of the server, and improves the automation and intelligence level of operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120434280A_ABST
    Figure CN120434280A_ABST
Patent Text Reader

Abstract

The invention, which relates to the technical field of information and communication, discloses an intelligent operation and maintenance monitoring system for a machine room, comprising a state monitoring module, an abnormity early warning module, a data acquisition module, a load evaluation module and a control module. The state monitoring module acquires temperature distribution state data information and vibration state data information in real time; the abnormity early warning module performs correlation analysis on the data, judges whether the server is normal or not and sends out an early warning signal; the data acquisition module acquires memory state data information and network transmission state data information of the server after receiving the early warning signal; the load evaluation module analyzes the load condition of the current server in combination with the trained load degree evaluation model; and the control module judges whether the server is in overload operation or not according to the current server load condition, generates a corresponding control instruction and executes the control instruction, so that stable operation and operation and maintenance efficiency of the server are guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information and communication technologies, and specifically to an intelligent operation and maintenance monitoring system for computer rooms. Background Art

[0002] With the rapid development of information and communication technologies, as the core infrastructure for information storage, processing, and transmission, the operation and management of data centers have become increasingly important. In this context, the efficient operation and maintenance of data centers play a crucial role in ensuring the continuity and stability of various services. And as an important part of data center management, computer room operation and maintenance monitoring is beneficial to detecting potential risks and taking corresponding measures, providing guarantee for the reliable operation of computer rooms.

[0003] However, currently, traditional computer room server operation and maintenance monitoring systems mainly rely on simple monitoring of independent devices and manual inspections. This mode usually can only obtain basic environmental data, lacking in-depth analysis of the running status, performance indicators, and potential anomalies of servers, resulting in delayed problem discovery and difficulty in early warning. First of all, the traditional monitoring system lacks the ability to process multi-dimensional data and perform intelligent analysis, and there is a data island effect among monitoring modules. Moreover, abnormal surface temperature distribution and vibration state changes of servers often indicate hardware overload. Once abnormal situations such as high temperature and vibration occur in the server, the traditional system is difficult to perceive and perform correlation analysis in a timely manner, easily leading to increased equipment overload and hardware loss. When the problem further develops, it may not only directly affect a significant decline in server performance and service interruption, but also cause system-level chain reactions, including serious consequences such as data loss and network paralysis, posing a threat to the overall stability and security of the computer room. Therefore, there is an urgent need for an intelligent and multi-dimensional correlation analysis computer room operation and maintenance monitoring system to remedy the defects of the existing technology and effectively address the above challenges. Summary of the Invention

[0004] In view of the deficiencies of the prior art, the present invention provides an intelligent operation and maintenance monitoring system for computer rooms, which solves the problems in the above background art.

[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: An intelligent operation and maintenance monitoring system for computer rooms includes a status monitoring module, an anomaly warning module, a data acquisition module, a load assessment module, and a control module;

[0006] The status monitoring module is used to monitor the surface temperature distribution and vibration state of computer room servers in real time, respectively obtaining temperature distribution state data information and vibration state data information;

[0007] The abnormal warning module is used to associate the temperature distribution state data information with the vibration state data information, analyze whether the current server is in a normal working state, and issue a warning signal;

[0008] The data acquisition module is used to collect the memory state data information of the current server after receiving the warning signal, and collect the network transmission state data information of the current server according to the processing situation of the received network transmission data by the current server;

[0009] The load evaluation module is used to associate the memory state data information with the network transmission state data information, and analyze the current server load situation in combination with the trained load degree evaluation model;

[0010] The control module is used to judge whether the current server is in an overloaded operation state according to the current server load situation, generate corresponding level control instructions and execute them.

[0011] Preferably, the state monitoring module includes a deployment unit, a monitoring unit and a preprocessing unit;

[0012] The deployment unit is used to deploy an infrared thermal imaging camera device on the outer surface of the computer room server, and fix a high-sensitivity acceleration sensor at the bottom of the server, and connect the infrared thermal imaging camera device and the high-sensitivity acceleration sensor to the central monitoring platform by means of wired connection. The central monitoring platform is used to preprocess the received relevant data and then perform a storage operation;

[0013] The monitoring unit is used to use the infrared thermal imaging camera device deployed on the outer surface of the server, and in combination with the automatic control technology, dynamically adjust the shooting distance, focal length and angle according to the size and shape characteristics of the server surface, capture and shoot the orthographic projection of the surface temperature distribution of the server, obtain the surface temperature distribution heat map, and use the wavelet denoising technology to remove the noise from the obtained surface temperature distribution heat map. And in combination with the gamma correction image enhancement technology, adjust the brightness distribution and contrast of the surface temperature distribution heat map; by uniformly dividing the surface temperature distribution heat map, generate several groups of surface heat map sub-regions, and respectively mark the several groups of surface heat map sub-regions as the first heat map sub-region S1, the second heat map sub-region S2, the third heat map sub-region S3,..., the nth heat map sub-region Sn; use the surface temperature distribution heat map, and in combination with the image processing algorithm, identify and extract the temperature distribution state data information of the outer surface of the server, and the temperature distribution state data information includes the red standard area Shc and the orange standard area Shh of each group of surface heat map sub-regions;

[0014] A highly sensitive acceleration sensor is fixed at the bottom position of the server to monitor the vibration state of the server in real time, construct vibration signals at each monitoring time point within the monitoring period, perform frequency-domain analysis on the vibration signals through Fourier transform, and obtain vibration state data information. The vibration state data information includes the vibration frequency Pzd and amplitude Vzf at each monitoring time point within the monitoring period.

[0015] The preprocessing unit is used to preprocess the temperature distribution state data information and vibration state data information received by the central monitoring platform, including filling missing values, removing duplicate data, filtering outliers, time series alignment, and normalization processing.

[0016] Preferably, the abnormal warning module includes a heat distribution analysis unit, a vibration analysis unit, and a warning unit.

[0017] The heat distribution analysis unit is used to determine the area Sz of a single heat map sub-region according to several groups of heat map sub-regions, construct the heat area occupancy ratio Vzb of several groups of heat map sub-regions by associating the area Sz of a single heat map sub-region with the temperature distribution state data information. Taking the heat area occupancy ratio Vzb of the i-th group of heat map sub-regions as an example, it is obtained specifically through the following formula: i For example, it is obtained specifically through the following formula:

[0018]

[0019] In the formula, Shc i represents the red-marked area of the i-th group of surface heat map sub-regions, and Shh i represents the orange-marked area of the i-th group of surface heat map sub-regions.

[0020] According to the method of obtaining the heat area occupancy ratio Vzb of the i-th group of heat map sub-regions, the heat area occupancy ratios Vzb of several groups of heat map sub-regions are obtained respectively, and the mean value of the heat area occupancy ratio is calculated according to the statistical mean algorithm. i By comparing the heat area occupancy ratios Vzb of several groups of heat map sub-regions with the mean value and marking the heat area occupancy ratios Vzb of the corresponding heat map sub-regions that exceed the mean value a heat difference sub-region set is constructed, and a relevant heat distribution imbalance signal is generated according to the number of heat map sub-regions in the heat difference sub-region set. The specific process is as follows: a heat difference sub-region set is constructed, and a relevant heat distribution imbalance signal is generated according to the number of heat map sub-regions in the heat difference sub-region set. The specific process is as follows:

[0021] When the number of heat map sub-regions in the heat difference sub-region concentration exceeds 30% of the total number of heat map sub-regions in the surface temperature distribution heat map, a "severe heat distribution imbalance signal" is generated. When the number of heat map sub-regions in the heat difference sub-region concentration does not exceed 30% of the total number of heat map sub-regions in the surface temperature distribution heat map, a "mild heat distribution imbalance signal" is generated, and both the generated "severe heat distribution imbalance signal" and "mild heat distribution imbalance signal" are sent to the warning unit;

[0022] The related heat distribution imbalance signals include a "severe heat distribution imbalance signal" and a "mild heat distribution imbalance signal".

[0023] Preferably, the vibration analysis unit is used to identify the obtained vibration state data information, and by using the mean algorithm in statistics, the average vibration frequency within the monitoring period is obtained and the average amplitude

[0024] Feature extraction is performed on the vibration state data information, a vibration model is established, the vibration model is trained and verified, and after linear normalization processing, a vibration anomaly coefficient Xbd is obtained. The vibration anomaly coefficient Xbd is obtained through the following formula:

[0025]

[0026] In the formula, Pzd j represents the vibration frequency at the j-th monitoring time point, represents the average vibration frequency within the monitoring period, Vzf j represents the amplitude at the j-th monitoring time point, represents the average amplitude within the monitoring period, where j = 1, 2,..., m, and m represents the number of monitoring time points;

[0027] By presetting a vibration anomaly reference threshold Z and comparing and analyzing the vibration anomaly reference threshold Z with the vibration anomaly coefficient Xbd of the server, relevant vibration anomaly signals are generated. The specific process is as follows:

[0028] When the vibration anomaly coefficient Xbd is greater than the preset vibration anomaly reference threshold Z, a "severe vibration anomaly signal" is generated. When the vibration anomaly coefficient Xbd is less than or equal to the preset vibration anomaly reference threshold Z, a "mild vibration anomaly signal" is generated, and both the generated "severe vibration anomaly signal" and "mild vibration anomaly signal" are sent to the warning unit;

[0029] The related vibration anomaly signals include a "severe vibration anomaly signal" and a "mild vibration anomaly signal".

[0030] Preferably, the warning unit is used to analyze whether the current server is in a normal working state by correlating the received relevant heat distribution imbalance signal and the relevant vibration anomaly signal, and issue a warning signal. The specific process is as follows:

[0031] Based on the relevant heat distribution imbalance signal, a set X is established. The "mild heat distribution imbalance signal" is designated as element a1, and the "severe heat distribution imbalance signal" is designated as element a2. And element a1 ∈ set X, element a2 ∈ set X;

[0032] Based on the relevant vibration anomaly signal, a set Y is established. The "mild vibration anomaly signal" is designated as element b1, and the "severe vibration anomaly signal" is designated as element b2. And element b1 ∈ set Y, element b2 ∈ set Y;

[0033] The union of set X and set Y is processed. If X ∪ Y = {a1, b2} or {a2, b1} or {a2, b2}, it means that the current server is not in a normal working state, and a warning signal is sent outwards. If X ∪ Y = {a1, b1}, it means that the current server is in a normal working state, and no additional warning signal is sent outwards.

[0034] Preferably, after receiving the warning signal, the data acquisition module is used to monitor the operating memory status of the central processing unit inside the current server and the remaining disk memory space, and construct memory status data information. The memory status data information includes the operating memory occupancy rate Vyx and the remaining storage memory Dcc. And according to the processing situation of the received network transmission data by the current server, using a network traffic capture tool, the network transmission status information of the current server is collected, and network transmission status data information is constructed. The network transmission status data information includes the data transmission rate Vcs, the packet loss rate Vdb, and the read / write delay duration Tys;

[0035] Both the operating memory occupancy rate and the network transmission status data information are sent to the central monitoring platform for data storage.

[0036] Preferably, the load evaluation module includes a memory analysis unit, a transmission analysis unit, and an evaluation unit;

[0037] The memory analysis unit is used to correlate the operating memory occupancy rate Vyx of the central processing unit with the remaining storage memory Dcc of the disk based on the memory status data information stored in the central monitoring platform. After dimensionless processing, the memory redundancy coefficient Xnc is obtained. The memory redundancy coefficient Xnc is obtained through the following formula:

[0038]

[0039] In the formula, both α1 and α2 represent weight values, and A represents the first correction constant;

[0040] The transmission analysis unit is used to analyze the network transmission status data information stored in the central monitoring platform. After dimensionless processing, the transmission anomaly coefficient Xcs is obtained. The transmission anomaly coefficient Xcs is obtained through the following formula:

[0041]

[0042] In the formula, Vcs represents the data transmission rate, Vcs max represents the maximum bandwidth supported by the current server, Vdb represents the packet loss rate, Tys represents the read / write delay duration, and β1, β2, and β3 all represent weight values.

[0043] Preferably, the evaluation unit is used to build an initial model using convolutional neural network technology, and train and test the initial model with the memory status data information and network transmission status data information. Then, the trained initial model is used as the relevant load association model, and the feature information in the relevant load association model is obtained respectively. The obtained feature information is used to train and test the relevant load association model, and a warning signal is sent in combination. The trained relevant load association model is used as the load level evaluation model;

[0044] The memory status data information is associated with the network transmission status data information. After dimensionless processing, combined with the trained load level evaluation model, the load comprehensive evaluation index Zfz is obtained by fitting. The load comprehensive evaluation index Zfz is obtained through the following formula;

[0045]

[0046] In the formula, Xnc represents the memory redundancy coefficient, Xcs represents the transmission anomaly coefficient, γ1 and γ2 respectively represent the weight values of the memory redundancy coefficient Xnc and the transmission anomaly coefficient Xcs, and B represents the second correction constant.

[0047] Preferably, the control module includes a degree comparison unit and a load control unit;

[0048] The degree comparison unit is used to preset an evaluation threshold P based on the historical data in the central monitoring platform, compare and analyze the load comprehensive evaluation index Zfz with the preset evaluation threshold P to determine whether the current server is in an overloaded operation state, and generate a corresponding level control instruction. The specific content is as follows:

[0049] If the load comprehensive evaluation index Zfz ≥ the evaluation threshold P, it means that the current server is in an overloaded operation state, and a first-level control instruction is generated and sent to the load control unit;

[0050] When the load comprehensive evaluation index Zfz < the evaluation threshold P, it indicates that the current server is not in an overloaded operation state, and a secondary control instruction is generated and sent to the load control unit.

[0051] Preferably, the load control unit is used to receive the primary control instruction and the secondary control instruction generated in the degree comparison unit, and execute the corresponding control operations. The specific content is as follows:

[0052] After receiving the primary control instruction, the execution content is: According to the current server being in an overloaded operation state, adjust the rotation speed of the server cooling fan; automatically trigger the server load balancing function, and enable the standby server device and the cloud server, dynamically migrate high-load tasks to the standby server device and the cloud server for processing, and at the same time limit the usage rights of high-resource-consuming applications; generate an overloaded operation report and an overloaded alarm, and send them to the operation and maintenance personnel; start the log recording function of the central monitoring platform, record the current server operation status information, and provide a reference for subsequent analysis and optimization;

[0053] After receiving the secondary control instruction, the execution content is: According to the current server not being in an overloaded operation state, continue to monitor the surface temperature and vibration state of the current server, as well as the memory occupancy and network transmission situation of the server; regularly generate a load status report, and at the same time store the load status report in the central monitoring platform, providing a reference basis for the optimization of the subsequent load degree evaluation model and system adjustment.

[0054] The present invention provides a computer room intelligent operation and maintenance monitoring system, which has the following beneficial effects:

[0055] (1) By integrating a status monitoring module, an anomaly warning module, a data acquisition module, a load assessment module, and a control module, a highly integrated and intelligent operation and maintenance monitoring system is constructed, achieving all-round perception and intelligent management of the operating status of the computer room servers. The system innovatively combines an infrared thermal imaging camera device and a high-sensitivity acceleration sensor, and through advanced algorithms such as gamma correction, wavelet denoising, and Fourier transform, realizes precise monitoring and data extraction of the surface temperature distribution and vibration status of the servers. Through multi-dimensional data correlation analysis, combining the temperature and vibration status, the working status of the servers is judged in real time and a warning signal is generated, improving the sensitivity and accuracy of anomaly detection; by further expanding the monitoring scope to cover the real-time acquisition and storage of the server memory and network transmission status, providing multi-dimensional data support for load assessment; by introducing convolutional neural network technology, deep learning analysis is carried out on the memory status and network transmission status, and a load degree assessment model is constructed. By calculating the comprehensive load assessment index Zfz, a quantitative basis is provided for the operating status of the servers; finally, according to the evaluation results, corresponding level control instructions are intelligently generated. When the servers are in an overloaded state, emergency measures such as load balancing strategies, task migration, and heat dissipation regulation can be quickly triggered, and an alarm is sent to the operation and maintenance personnel; when not in an overloaded state, continuous monitoring and optimization are carried out to ensure the stability and operation efficiency of the servers; overall, the present invention improves the automation and intelligence level of computer room operation and maintenance, effectively solves the problems of insufficient multi-dimensional data analysis, response lag, and poor correlation in traditional systems, guarantees the business continuity, security, and high efficiency of the computer room, and provides a systematic solution for modern computer room operation and maintenance management.

[0056] (2) Through technological innovation, the status monitoring module and the anomaly warning module achieve precise monitoring and early warning of the temperature and vibration status of the servers, effectively overcoming the limitations of single-dimensional data monitoring and manual inspection in traditional monitoring systems; in the status monitoring module, through the infrared thermal imaging camera device, combined with gamma correction and wavelet denoising technologies, a high-precision surface temperature distribution heat map is generated, and the server surface is evenly divided and marked to obtain temperature distribution status data information, providing a more targeted analysis basis for anomaly monitoring; at the same time, by deploying a high-sensitivity acceleration sensor, frequency domain analysis of the vibration signal is carried out using Fourier transform to obtain vibration status data information, which can accurately capture early signs of hardware failures; the anomaly warning module conducts multi-dimensional correlation analysis on the two types of data. Through the classification analysis of relevant heat distribution imbalance signals and relevant vibration anomaly signals, different failure types of the servers can be accurately distinguished, and a warning signal is generated according to the severity; this design effectively solves the problem of difficult correlation of multiple anomaly factors in traditional monitoring, shortens the problem monitoring time, provides early warning for subsequent analysis, and thus improves the risk control ability of the computer room.

[0057] (3) By introducing deep learning technology and dimensionless processing methods, it provides scientificity and accuracy for the load status analysis of the server. In the load evaluation module, by combining multiple key indicators such as the running memory occupancy rate Vyx, the remaining storage memory Dcc, the data transmission rate Vcs, the packet loss rate Vdb, and the read / write latency duration Tys, the memory redundancy coefficient Xnc and the transmission anomaly coefficient Xcs are constructed. And a load degree evaluation model trained by convolutional neural network technology is used to generate a comprehensive load evaluation index Zfz. The comprehensive load evaluation index Zfz provides an accurate quantitative basis for the server load status by comparing with a preset evaluation threshold P. In terms of intelligent control, corresponding control instructions are generated and executed according to different levels of the load status. When the server is in an overloaded running state, the load balancing function is automatically triggered, high-load tasks are migrated to the standby server and the cloud server, and at the same time, the usage rights of high-resource-consuming applications are restricted to relieve the server pressure. For the normal load status, the status monitoring continues, and a load status report is generated regularly. This refined load management mechanism not only improves the resource utilization rate but also avoids the risks of hardware loss and service interruption caused by overload, providing a reliable guarantee for the efficiency and security of the intelligent operation and maintenance of the computer room. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 It is a block diagram of an intelligent operation and maintenance monitoring system for a computer room according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0059] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0060] Embodiment 1

[0061] Please refer to Figure 1 , the present invention provides an intelligent operation and maintenance monitoring system for a computer room, including a status monitoring module, an anomaly warning module, a data acquisition module, a load evaluation module, and a control module;

[0062] The status monitoring module is used to monitor the surface temperature distribution and vibration state of the computer room server in real time, and obtain the temperature distribution state data information and vibration state data information respectively;

[0063] The anomaly warning module is used to associate the temperature distribution state data information with the vibration state data information, analyze whether the current server is in a normal working state, and issue a warning signal;

[0064] The data acquisition module is used to collect the memory status data information of the current server after receiving the warning signal, and collect the network transmission status data information of the current server according to the processing situation of the received network transmission data by the current server;

[0065] The load evaluation module is used to associate the memory status data information with the network transmission status data information, and analyze the current server load situation in combination with the trained load degree evaluation model;

[0066] The control module is used to judge whether the current server is in an overloaded operation state according to the current server load situation, and generate corresponding level control instructions and execute them.

[0067] In this embodiment, through the multi-module collaboration of status monitoring, abnormal warning, data acquisition, load evaluation and intelligent control, the deficiencies of the traditional computer room monitoring system in multi-dimensional data processing ability and intelligent analysis are effectively solved; firstly, while real-time obtaining the temperature distribution and vibration state, combined with the multi-dimensional data correlation analysis technology, the recognition accuracy and warning ability of potential server faults are improved, thus compensating for the problems of response lag and data island effect in the traditional monitoring system; in addition, through the introduction of in-depth monitoring of memory status and network transmission status and the trained load comprehensive evaluation model, the accurate quantitative analysis of the server operation state is realized, making the discovery of the overloaded state faster and more reliable; the dynamic task migration, load balancing and automatic response mechanism of the control module can not only effectively alleviate the hardware loss and service interruption risks caused by overloading, but also ensure the resource optimization and continuous monitoring of the server under the normal load state; this systematic and intelligent design improves the efficiency and security of computer room operation and maintenance, effectively overcomes the shortcomings of the traditional mode in data acquisition, analysis and decision-making, and provides strong technical support for the stable operation and business continuity of modern computer room operation and maintenance.

[0068] Embodiment 2

[0069] Please refer to Figure 1 Specifically, the status monitoring module includes a deployment unit, a monitoring unit and a preprocessing unit;

[0070] The deployment unit is used to deploy an infrared thermal imaging camera device on the outer surface of the computer room server, and fix a high-sensitivity acceleration sensor at the bottom of the server, and connect the infrared thermal imaging camera device and the high-sensitivity acceleration sensor to the central monitoring platform by means of wired connection. The central monitoring platform is used to preprocess the received relevant data and then perform a storage operation;

[0071] The monitoring unit is used to utilize the infrared thermal imaging camera devices deployed on the outer surface of the server, and in combination with automated control technology, dynamically adjust the shooting distance, focal length, and angle according to the size and shape characteristics of the server surface, capture and shoot the orthographic projection of the surface temperature distribution of the server, obtain the surface temperature distribution heat map, remove noise from the obtained surface temperature distribution heat map through wavelet denoising technology, and adjust the brightness distribution and contrast of the surface temperature distribution heat map in combination with gamma correction image enhancement technology; by uniformly dividing the surface temperature distribution heat map, several groups of surface heat map sub-regions are generated, and the several groups of surface heat map sub-regions are respectively marked as the first heat map sub-region S1, the second heat map sub-region S2, the third heat map sub-region S3,..., the nth heat map sub-region Sn; utilize the surface temperature distribution heat map, and in combination with image processing algorithms, identify and extract the temperature distribution state data information on the outer surface of the server, and the temperature distribution state data information includes the red-labeled area Shc and the orange-labeled area Shh of each group of surface heat map sub-regions;

[0072] It should be noted that the red-labeled area Shc and the orange-labeled area Shh are key parameters in the server surface temperature distribution heat map, which are used to represent the area size of the higher temperature regions; among them, the red-labeled area Shc represents the area of the region in the heat map where the temperature exceeds the upper limit of the hardware safe operating temperature; the orange-labeled area Shh represents the area of the region where the temperature is in the medium-high temperature range, close to the high temperature threshold but not exceeding the standard.

[0073] A high-sensitivity acceleration sensor is fixed at the bottom of the server to monitor the vibration state of the server in real time, construct vibration signals at each monitoring time point during the monitoring period, perform frequency domain analysis on the vibration signals through Fourier transform, and obtain vibration state data information, and the vibration state data information includes the vibration frequency Pzd and amplitude Vzf at each monitoring time point during the monitoring period;

[0074] The preprocessing unit is used to preprocess the temperature distribution state data information and vibration state data information received by the central monitoring platform, including filling missing values, removing duplicate data, filtering outliers, time series alignment, and normalization processing.

[0075] In this embodiment, through the collaborative design of the deployment unit, monitoring unit, and preprocessing unit, the monitoring ability of the operation status of the computer room servers is optimized, effectively solving the problems of single monitoring, insufficient data analysis ability, and lagging anomaly detection in the traditional system. First, the deployment unit realizes multi-dimensional real-time data collection of temperature distribution and vibration status by deploying an infrared thermal imaging device on the outer surface of the server and fixing a high-sensitivity acceleration sensor at the bottom. The infrared thermal imaging device combined with the automatic control technology can dynamically adjust the shooting distance, focal length, and angle according to the size and shape of the server, generating a high-precision surface temperature distribution thermal map. The background interference in the surface temperature distribution thermal map is removed through wavelet denoising technology, and combined with the gamma correction image enhancement technology, the brightness and contrast are further optimized to ensure the high accuracy and readability of the temperature data. The monitoring unit also evenly divides the surface temperature distribution thermal map, refines it into several groups of surface thermal map sub-regions, and extracts the temperature distribution data of each surface thermal map sub-region, including the red-labeled area Shc and the orange-labeled area Shh, providing fine-grained basic data for discovering abnormal heat distribution. At the same time, the high-sensitivity acceleration sensor is used for vibration status monitoring, and the vibration signal is analyzed in the frequency domain through Fourier transform to extract key characteristic parameters, including the vibration frequency Pzd and amplitude Vzf at each monitoring time point, accurately capturing the abnormal characteristics of the server. The preprocessing unit further cleans and standardizes the collected temperature and vibration data, including filling missing values, removing duplicate data, filtering outliers, time series alignment, and normalization processing, ensuring the high quality and consistency of the data, providing a reliable guarantee for the analysis of subsequent modules. This multi-dimensional and refined monitoring design improves the intelligent level of computer room operation and maintenance, not only effectively solving the problem of insufficient complex anomaly detection ability in the traditional system, but also providing reliable data support for the early warning of potential risks.

[0076] Embodiment 3

[0077] Please refer to Figure 1 , specifically: The anomaly warning module includes a heat distribution analysis unit, a vibration analysis unit, and a warning unit;

[0078] The heat distribution analysis unit is used to determine the area Sz of a single thermal map sub-region according to several groups of thermal map sub-regions, and by associating the area Sz of a single thermal map sub-region with the temperature distribution status data information, construct the heat area occupancy ratio Vzb of several groups of thermal map sub-regions. Taking the heat area occupancy ratio Vzb of the i-th group of thermal map sub-regions as an example, it is specifically obtained through the following formula: i For example, it is obtained through the following formula:

[0079]

[0080] In the formula, Shc iThe red - marked area representing the i - th group of surface heat - map sub - regions, Shh i The orange - marked area representing the i - th group of surface heat - map sub - regions;

[0081] According to the method of obtaining the ratio Vzb of the heat area of the i - th group of heat - map sub - regions, i obtain the ratios Vzb of the heat areas of several groups of heat - map sub - regions respectively, and calculate the average value of the ratio of the heat area according to the statistical mean - value algorithm. By comparing the ratios Vzb of the heat areas of several groups of heat - map sub - regions with the average value make a size comparison, mark the ratios Vzb of the heat areas of the corresponding heat - map sub - regions that exceed the average value to construct a set of heat - difference sub - regions, and generate relevant heat - distribution imbalance signals according to the number of heat - map sub - regions in the set of heat - difference sub - regions. The specific process is as follows:

[0082] When the number of heat - map sub - regions in the set of heat - difference sub - regions exceeds 30% of the total number of heat - map sub - regions in the surface - temperature - distribution heat - map, generate a "severe heat - distribution imbalance signal". When the number of heat - map sub - regions in the set of heat - difference sub - regions does not exceed 30% of the total number of heat - map sub - regions in the surface - temperature - distribution heat - map, generate a "mild heat - distribution imbalance signal", and send both the generated "severe heat - distribution imbalance signal" and "mild heat - distribution imbalance signal" to the warning unit;

[0083] The relevant heat - distribution imbalance signals include a "severe heat - distribution imbalance signal" and a "mild heat - distribution imbalance signal".

[0084] Specifically: The vibration analysis unit is used to identify the obtained vibration state data information, and obtain the average vibration frequency within the monitoring period through the mean - value algorithm in statistics and the average amplitude

[0085] Extract the features of the vibration state data information, establish a vibration model, train and verify the vibration model, and after linear normalization processing, obtain the vibration anomaly coefficient Xbd. The vibration anomaly coefficient Xbd is obtained through the following formula:

[0086]

[0087] In the formula, Pzd j represents the vibration frequency at the j - th monitoring time point, represents the average vibration frequency within the monitoring period, Vzf j represents the amplitude at the j - th monitoring time point, It is expressed as the average amplitude within the monitoring period, where j = 1, 2,..., m, and m represents the number of monitoring time points;

[0088] By presetting the vibration anomaly reference threshold Z and comparing and analyzing the vibration anomaly reference threshold Z with the vibration anomaly coefficient Xbd of the server, relevant vibration anomaly signals are generated. The specific process is as follows:

[0089] When the vibration anomaly coefficient Xbd is greater than the preset vibration anomaly reference threshold Z, a "severe vibration anomaly signal" is generated. When the vibration anomaly coefficient Xbd is less than or equal to the preset vibration anomaly reference threshold Z, a "mild vibration anomaly signal" is generated, and both the generated "severe vibration anomaly signal" and "mild vibration anomaly signal" are sent to the warning unit;

[0090] The relevant vibration anomaly signals include a "severe vibration anomaly signal" and a "mild vibration anomaly signal".

[0091] Specifically, the warning unit is used to analyze whether the current server is in a normal working state by associating the received relevant heat distribution imbalance signals and relevant vibration anomaly signals, and issue a warning signal. The specific process is as follows:

[0092] Based on the relevant heat distribution imbalance signals, a set X is established. The "mild heat distribution imbalance signal" is labeled as element a1, and the "severe heat distribution imbalance signal" is labeled as element a2, and element a1 ∈ set X, element a2 ∈ set X;

[0093] Based on the relevant vibration anomaly signals, a set Y is established. The "mild vibration anomaly signal" is labeled as element b1, and the "severe vibration anomaly signal" is labeled as element b2, and element b1 ∈ set Y, element b2 ∈ set Y;

[0094] The union of set X and set Y is processed. If X ∪ Y = {a1, b2} or {a2, b1} or {a2, b2}, it means that the current server is not in a normal working state, and a warning signal is sent outwards. If X ∪ Y = {a1, b1}, it means that the current server is in a normal working state, and no additional warning signal is sent outwards.

[0095] In this embodiment, through the collaborative design of the heat distribution analysis unit, vibration analysis unit and warning unit, the abnormal detection ability of the server operation state is optimized, effectively solving the problems of insufficient multi-dimensional data analysis and response lag in the traditional computer room monitoring system; the heat distribution analysis unit constructs the heat area occupancy ratio Vzb of the heat map sub-region by evenly dividing the heat map of the server surface temperature distribution and measuring the area of the sub-region, and combines the statistical mean algorithm to identify the values exceeding its mean The heat map sub-regions are marked to construct a set of sub-regions with heat differences. By comparing the proportion of the number of heat map sub-regions, "mild heat distribution imbalance signal" or "severe heat distribution imbalance signal" is dynamically generated. This mechanism can quickly capture local temperature anomalies and provide accurate early warning data for potential hardware overheating, effectively compensating for the deficiency that traditional monitoring systems are difficult to perform fine-grained heat distribution analysis. At the same time, the vibration analysis unit performs frequency domain analysis on the collected vibration signals through Fourier transform, extracts the core characteristic parameters of the vibration state, including the average vibration frequency during the monitoring period and the average amplitude Based on these characteristic parameters, a vibration model is established. After linearly normalizing the vibration data, the vibration anomaly coefficient Xbd is calculated. By comparing the vibration anomaly coefficient Xbd with the preset vibration anomaly reference threshold Z, the severity of vibration anomalies can be quickly judged, and "mild vibration anomaly signal" and "severe vibration anomaly signal" are generated. This high-precision vibration analysis method can effectively identify potential abnormal state problems and effectively overcome the limitation that traditional monitoring systems have insufficient monitoring of vibration states. The warning unit forms an abnormal judgment mechanism with multi-dimensional signal linkage through the correlation analysis of relevant heat distribution imbalance signals and relevant vibration anomaly signals. When the combined state of relevant heat distribution imbalance signals and relevant vibration anomaly signals indicates that the server is in an abnormal working state, a warning signal is immediately generated, improving the accuracy of abnormal monitoring. This method of multi-dimensional data correlation analysis effectively avoids false alarms and missed alarms caused by single-dimensional analysis in traditional monitoring systems and provides more accurate monitoring and warning capabilities for the server operation state. In addition, by refining data processing and intelligent signal analysis, the passive abnormal state processing mode is transformed into an active abnormal warning mode, thereby reducing the risks brought by hardware failures to business continuity and system stability. Through the accurate classification and rapid response of relevant abnormal signals, it can effectively prevent hardware losses and performance degradation caused by overload, and at the same time reduce the troubleshooting time and workload of operation and maintenance personnel. This modular and multi-dimensional abnormal detection and warning mechanism not only surpasses traditional monitoring systems in terms of response speed and abnormal recognition accuracy, but also effectively overcomes the technical shortcomings of data islands and insufficient analysis dimensions in the traditional mode, providing new technical support for the intelligence and efficiency of computer room operation and maintenance.

[0096] Embodiment 4

[0097] Please refer to Figure 1, specifically: after receiving the warning signal, the data acquisition module monitors the operating memory status of the central processing unit inside the current server and the remaining disk memory space, constructs memory status data information, and the memory status data information includes the operating memory occupancy rate Vyx and the remaining storage memory Dcc. According to the processing situation of the received network transmission data by the current server, using a network traffic capture tool, it collects the network transmission status information of the current server and constructs network transmission status data information, and the network transmission status data information includes the data transmission rate Vcs, the packet loss rate Vdb, and the read / write delay duration Tys;

[0098] It should be noted that the operating memory occupancy rate Vyx represents the proportion of the memory used by the current central processing unit in the total memory capacity, and is obtained through the memory monitoring interface provided by the operating system; the remaining storage memory Dcc represents the unused capacity in the disk storage space and is detected in real time through the disk management interface of the storage device; the data transmission rate Vcs represents the amount of data received or sent by the server per unit time and is obtained by statistics through the network interface of the operating system; the packet loss rate Vdb represents the proportion of packets lost during the transmission process and is calculated by analyzing the transmission control protocol in the network protocol stack; the read / write delay duration Tys represents the delay time between the storage device and the network transmission of the data and is measured by a performance analysis tool; regularly upload the collected relevant data to the central monitoring platform, record and process it to form historical data, providing basic data support for subsequent analysis;

[0099] Send both the operating memory occupancy rate and the network transmission status data information to the central monitoring platform for data storage.

[0100] In this embodiment, by monitoring the internal memory status and network transmission status of the server in real time, the system's perception ability of the server performance status is improved, effectively solving the problems of single-dimensional internal status monitoring and insufficient dynamic acquisition in traditional monitoring systems; by collecting the running memory occupancy rate Vyx and the remaining storage memory Dcc, the usage of the server memory and storage pressure can be accurately captured, providing key data support for the optimized management of memory resources. At the same time, by using a network traffic capture tool to collect the data transmission rate Vcs, packet loss rate Vdb, and read / write latency duration Tys in real time, the system can master the actual performance status of the server in network transmission, including key indicators such as traffic bottlenecks, latency, and data transmission reliability. This design makes up for the deficiency of network performance monitoring in traditional systems. Especially when dealing with high-traffic loads, potential anomalies can be dynamically identified and quickly responded to, and the running memory occupancy rate and network transmission status data information are sent to the central monitoring platform for data storage, providing a high-quality data foundation for subsequent load assessment and intelligent control. This multi-dimensional data collection mechanism not only surpasses the traditional mode in the accuracy and real-time performance of performance monitoring but also provides more reliable data support for the subsequent server running load by constructing a complete server running status system, reducing the risks of performance degradation and service interruption caused by the failure to detect memory and network transmission anomalies in a timely manner, and ensuring the efficient operation of the server and the business continuity of the computer room.

[0101] Embodiment 5

[0102] Please refer to Figure 1 , specifically: The load assessment module includes a memory analysis unit, a transmission analysis unit, and an assessment unit;

[0103] The memory analysis unit is used to associate the running memory occupancy rate Vyx of the central processing unit with the remaining storage memory Dcc of the disk according to the memory status data information stored in the central monitoring platform. After dimensionless processing, the memory redundancy coefficient Xnc is obtained. The memory redundancy coefficient Xnc is obtained through the following formula:

[0104]

[0105] In the formula, both α1 and α2 represent weight values, and A represents a first correction constant;

[0106] The transmission analysis unit is used to analyze the network transmission status data information stored in the central monitoring platform. After dimensionless processing, the transmission anomaly coefficient Xcs is obtained. The transmission anomaly coefficient Xcs is obtained through the following formula:

[0107]

[0108] Where, Vcs represents the data transmission rate, Vcs max represents the maximum bandwidth supported by the current server, Vdb represents the packet loss rate, Tys represents the read / write delay duration, and β1, β2, and β3 all represent weight values.

[0109] Specifically, the evaluation unit is used to construct an initial model using convolutional neural network technology, and train and test the initial model with memory state data information and network transmission state data information, and use the trained initial model as a related load association model, respectively obtain the feature information in the related load association model, and train and test the related load association model with the obtained feature information, combine to issue a warning signal, and use the trained related load association model as a load degree evaluation model;

[0110] Associate the memory state data information with the network transmission state data information, after dimensionless processing, combine with the trained load degree evaluation model, and fit to obtain the load comprehensive evaluation index Zfz. The load comprehensive evaluation index Zfz is obtained through the following formula;

[0111]

[0112] Where, Xnc represents the memory redundancy coefficient, Xcs represents the transmission anomaly coefficient, γ1 and γ2 respectively represent the weight values of the memory redundancy coefficient Xnc and the transmission anomaly coefficient Xcs, and B represents the second correction constant.

[0113] In this embodiment, through the collaborative work of the memory analysis unit, the transmission analysis unit and the evaluation unit, the accurate quantification and in-depth analysis of the server load status are achieved, effectively solving the problems of insufficient multi-dimensional load correlation analysis ability and single evaluation method in the traditional system; the memory analysis unit correlates the collected running memory occupancy rate Vyx with the remaining storage memory of the disk Dcc, and combines the dimensionless processing method to calculate the memory redundancy coefficient Xnc, which can intuitively reflect the pressure status of the server memory usage and provide important data support for optimizing resource allocation; the transmission analysis unit is based on the network performance indicators of data transmission rate Vcs, packet loss rate Vdb and read / write delay duration Tys, and combines the dimensionless processing technology to generate the transmission anomaly coefficient Xcs, systematically evaluating the network transmission efficiency and stability of the server; the evaluation unit further correlates the memory status data information with the network transmission status data information, and after dimensionless processing, combines the trained load degree evaluation model to fit and obtain the comprehensive load evaluation index Zfz. The comprehensive load evaluation index Zfz reflects the overall load status of the server and provides a scientific basis for judging whether the server is in an overloaded operation state; this multi-dimensional and intelligent load evaluation method not only effectively overcomes the defect of inaccurate evaluation of complex load problems in the traditional system, but also improves the real-time and reliability of load evaluation through the correlation analysis of abnormal signals. Through the quantitative analysis of the comprehensive load evaluation index Zfz, the potential overload risk of the server can be quickly identified, providing high-quality data support for the response decision of the subsequent control module, thus effectively avoiding the problems of server performance degradation and service interruption caused by equipment overload and increased hardware loss, and improving the operation efficiency and stability of the computer room.

[0114] Embodiment 6

[0115] Please refer to Figure 1 , specifically: the control module includes a degree comparison unit and a load control unit;

[0116] The degree comparison unit is used to preset the evaluation threshold P according to the historical data in the central monitoring platform, compare and analyze the comprehensive load evaluation index Zfz with the preset evaluation threshold P to judge whether the current server is in an overloaded operation state, and generate corresponding level control instructions, the specific content is as follows:

[0117] If the comprehensive load evaluation index Zfz ≥ the evaluation threshold P, it means that the current server is in an overloaded operation state, and a first-level control instruction is generated and sent to the load control unit;

[0118] If the comprehensive load evaluation index Zfz < the evaluation threshold P, it means that the current server is not in an overloaded operation state, and a second-level control instruction is generated and sent to the load control unit.

[0119] Specifically, the load control unit is used to receive the primary control instruction and the secondary control instruction generated by the degree comparison unit and execute the corresponding control operations. The specific content is as follows:

[0120] After receiving the primary control instruction, the execution content is as follows: According to the current server being in an overloaded operation state, adjust the rotation speed of the server cooling fan; automatically trigger the server load balancing function and enable the standby server device and cloud server, dynamically migrate high-load tasks to the standby server device and cloud server for processing, and at the same time limit the usage rights of high-resource-consuming applications; generate an overloaded operation report and an overloaded alarm, and send them to the operation and maintenance personnel; start the log recording function of the central monitoring platform, record the current server operation status information, and provide a reference for subsequent analysis and optimization;

[0121] After receiving the secondary control instruction, the execution content is as follows: According to the current server not being in an overloaded operation state, continue to monitor the surface temperature and vibration state of the current server, as well as continuously monitor the memory occupancy and network transmission situation of the server; regularly generate a load status report, and at the same time store the load status report in the central monitoring platform, providing a reference basis for the optimization of the subsequent load degree evaluation model and system adjustment.

[0122] In this embodiment, through the collaborative design of the degree comparison unit and the load control unit, the intelligent response and dynamic regulation of the server load status are achieved, effectively solving the deficiencies of the traditional computer room monitoring system in load management and emergency response capabilities; the degree comparison unit, based on the historical data of the central monitoring platform, judges whether the current server is in an overloaded operation state by comparing and analyzing the load comprehensive evaluation index Zfz with the preset evaluation threshold P, and generates corresponding level control instructions, providing a scientific basis for the intelligent response to different load states; when the server is in an overloaded operation state, after receiving the first-level control instruction, the load control unit quickly executes a series of emergency measures including adjusting the rotation speed of the cooling fan, triggering the load balancing function, and dynamic task migration to reduce the server load pressure and prevent hardware loss. At the same time, by restricting the usage rights of high-resource-consuming applications and sending overloaded alarms and operation reports to the operation and maintenance personnel, the risks of hardware failures and service interruptions caused by overload can be effectively prevented; when the server is in an overloaded operation state, after receiving the second-level control instruction, the load control unit continues to perform real-time status monitoring and regularly generates load status reports to ensure the continuity of the server operation and the scientific nature of the optimization adjustment; this multi-level and dynamic control strategy can not only quickly relieve the operation pressure under high load, but also provide support for subsequent optimization through continuous monitoring and data accumulation, compensating for the lag and lack of flexibility in load regulation of the traditional system; through the intelligent judgment and response mechanism, the operation efficiency and safety of the computer room are improved, and the risks of hardware loss and service interruption caused by abnormal load are reduced.

[0123] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent operation and maintenance monitoring system for a computer room, characterized by: It includes status monitoring module, abnormal warning module, data acquisition module, load assessment module and control module; The state monitoring module is used to monitor the surface temperature distribution and vibration status of the server in the computer room in real time, and obtain temperature distribution status data information and vibration status data information respectively; The abnormal warning module is used to associate the temperature distribution state data information with the vibration state data information, analyze whether the current server is in a normal working state, and issue a warning signal; The data acquisition module is used to collect the memory status data information of the current server after receiving the warning signal, and collect the network transmission status data information of the current server according to the current server's processing of the received network transmission data; The load evaluation module is used to associate the memory status data information with the network transmission status data information, and analyze the current server load situation in combination with the trained load degree evaluation model; The control module is used to determine whether the current server is in an overloaded operating state based on the current server load situation, generate corresponding level control instructions and execute them.

2. The intelligent operation and maintenance monitoring system for a computer room according to claim 1, characterized in that: The state monitoring module includes a deployment unit, a monitoring unit and a pre-processing unit; The deployment unit is used to deploy infrared thermal imaging equipment on the outer surface of the server in the computer room, and fix a high-sensitivity acceleration sensor at the bottom of the server, and connect the infrared thermal imaging equipment and the high-sensitivity acceleration sensor to the central monitoring platform through a wired connection. The central monitoring platform is used to pre-process the received relevant data and then perform storage operations; The monitoring unit is used to utilize infrared thermal imaging equipment deployed on the outer surface of the server, and in combination with automatic control technology, dynamically adjust the shooting distance, focal length and angle according to the size and shape characteristics of the server surface, capture and shoot the surface temperature distribution of the server by orthographic projection, and obtain a surface temperature distribution heat map. The obtained surface temperature distribution heat map is de-noised by wavelet denoising technology, and the brightness distribution and contrast of the surface temperature distribution heat map are adjusted by combining gamma correction image enhancement technology; the surface temperature distribution heat map is evenly divided to generate several groups of surface heat map sub-regions, and the several groups of surface heat map sub-regions are respectively marked as heat map sub-region No. 1 S1, heat map sub-region No. 2 S2, heat map sub-region No. 3 S3, ..., heat map sub-region No. n Sn; the surface temperature distribution heat map is used, in combination with an image processing algorithm, to identify and extract temperature distribution status data information of the outer surface of the server, wherein the temperature distribution status data information includes the red marked area Shc and the orange marked area Shh of each group of surface heat map sub-regions; Using a high-sensitivity acceleration sensor fixed at the bottom of the server, the vibration state of the server is monitored in real time, and vibration signals at each monitoring time point within the monitoring period are constructed. The vibration signals are analyzed in the frequency domain through Fourier transform to obtain vibration state data information. The vibration state data information includes the vibration frequency Pzd and amplitude Vzf at each monitoring time point within the monitoring period; The preprocessing unit is used to preprocess the temperature distribution state data information and vibration state data information received by the central monitoring platform, including filling missing values, removing duplicate data, filtering outliers, time series alignment and normalization processing.

3. The intelligent operation and maintenance monitoring system for a computer room according to claim 2, characterized in that: The abnormal warning module includes a heat distribution analysis unit, a vibration analysis unit and a warning unit; The heat distribution analysis unit is used to determine the area Sz of a single heat map sub-region based on the multiple groups of heat map sub-regions, and to construct the heat area ratio Vzb of the multiple groups of heat map sub-regions by associating the single heat map sub-region area Sz with the temperature distribution state data information; According to the heat area ratio Vzb of several groups of heat map sub-regions constructed, and combined with the statistical mean algorithm, the mean of the heat area ratio is calculated ; By comparing the heat area ratio Vzb of several groups of heat map sub-regions with the mean Comparing the size will exceed the mean The heat area ratio Vzb of the corresponding heat map sub-region is marked, and a heat difference sub-region set is constructed. According to the number of heat map sub-regions in the heat difference sub-region set, a related heat distribution imbalance signal is generated. The specific process is as follows: When the number of heat map sub-regions in the heat difference sub-region concentration exceeds 30% of the total number of heat map sub-regions in the surface temperature distribution heat map, a "heavy heat distribution imbalance signal" is generated. When the number of heat map sub-regions in the heat difference sub-region concentration does not exceed 30% of the total number of heat map sub-regions in the surface temperature distribution heat map, a "mild heat distribution imbalance signal" is generated. Both the generated "heavy heat distribution imbalance signal" and "mild heat distribution imbalance signal" are sent to the early warning unit. The relevant heat distribution imbalance signal includes a "heat distribution imbalance severe signal" and a "heat distribution imbalance mild signal".

4. The intelligent operation and maintenance monitoring system for a computer room according to claim 3, characterized in that: The vibration analysis unit is used to identify the acquired vibration state data information and obtain the average vibration frequency during the monitoring period through the statistical averaging algorithm. and average amplitude ; Extracting features from the vibration state data information, establishing a vibration model, training and validating the vibration model, and performing linear normalization processing to obtain a vibration abnormality coefficient Xbd; By presetting a vibration abnormality reference threshold Z and comparing and analyzing the vibration abnormality reference threshold Z with the vibration abnormality coefficient Xbd of the server, a relevant vibration abnormality signal is generated. The specific process is as follows: When the vibration abnormality coefficient Xbd is greater than the preset vibration abnormality reference threshold Z, a "severe vibration abnormality signal" is generated; when the vibration abnormality coefficient Xbd is less than or equal to the preset vibration abnormality reference threshold Z, a "mild vibration abnormality signal" is generated. Both the generated "severe vibration abnormality signal" and "mild vibration abnormality signal" are sent to the early warning unit; The relevant vibration abnormality signals include "severe vibration abnormality signals" and "mild vibration abnormality signals".

5. The intelligent operation and maintenance monitoring system for a computer room according to claim 4, characterized in that: The early warning unit is used to analyze whether the current server is in a normal working state by correlating the received related heat distribution imbalance signal and the related vibration abnormality signal, and issue an early warning signal. The specific process is as follows: Establish a set X based on the relevant heat distribution imbalance signal, mark the "heat distribution imbalance signal with mild intensity" as element a1, mark the "heat distribution imbalance signal with severe intensity" as element a2, and element a1∈set X, element a2∈set X; A set Y is established based on the relevant abnormal vibration signals, and the "mild abnormal vibration signal" is marked as element b1, and the "severe abnormal vibration signal" is marked as element b2, and element b1∈set Y, element b2∈set Y; Perform union processing on set X and set Y. If X∪Y={a1, b2} or {a2, b1} or {a2, b2}, it means that the current server is not in normal working condition, and an early warning signal is issued. If X∪Y={a1, b1}, it means that the current server is in normal working condition, and no additional early warning signal is issued.

6. The intelligent operation and maintenance monitoring system for a computer room according to claim 5, characterized in that: The data acquisition module is used to monitor the CPU running memory status and disk memory space remaining status of the current server after receiving the early warning signal, and construct memory status data information, wherein the memory status data information includes the running memory occupancy rate Vyx and the storage memory remaining amount Dcc, and based on the current server's processing of received network transmission data, use the network traffic capture tool to collect the network transmission status information of the current server, and construct network transmission status data information, wherein the network transmission status data information includes the data transmission rate Vcs, the packet loss rate Vdb and the read and write delay time Tys; The running memory occupancy rate and network transmission status data information are sent to the central monitoring platform for data storage.

7. The intelligent operation and maintenance monitoring system for a computer room according to claim 6, characterized in that: The load evaluation module includes a memory analysis unit, a transmission analysis unit and an evaluation unit; The memory analysis unit is used to associate the CPU's running memory occupancy rate Vyx with the disk's storage memory remaining capacity Dcc based on the memory status data information stored in the central monitoring platform, and obtain the memory redundancy coefficient Xnc after dimensionless processing; The transmission analysis unit is used to analyze the network transmission status data information stored in the central monitoring platform, and obtain the transmission anomaly coefficient Xcs after dimensionless processing.

8. The intelligent operation and maintenance monitoring system for a computer room according to claim 7, characterized in that: The evaluation unit is used to construct an initial model using convolutional neural network technology, and train and test the initial model using memory status data information and network transmission status data information, and use the trained initial model as a relevant load association model, respectively obtain feature information within the relevant load association model, and train and test the relevant load association model using the obtained feature information, and use the trained relevant load association model as a load degree evaluation model in combination with issuing an early warning signal; The memory status data information is associated with the network transmission status data information, and after dimensionless processing, combined with the trained load degree evaluation model, a load comprehensive evaluation index Zfz is obtained by fitting.

9. The intelligent operation and maintenance monitoring system for a computer room according to claim 8, characterized in that: The control module includes a degree comparison unit and a load control unit; The degree comparison unit is used to pre-set an evaluation threshold P based on historical data in the central monitoring platform, compare and analyze the load comprehensive evaluation index Zfz with the preset evaluation threshold P to determine whether the current server is in an overloaded operating state and generate a corresponding level control instruction. The specific contents are as follows: If the load comprehensive evaluation index Zfz ≥ the evaluation threshold P, it indicates that the current server is in an overloaded state, and a first-level control instruction is generated and sent to the load control unit; If the load comprehensive evaluation index Zfz is less than the evaluation threshold P, it indicates that the current server is not in an overloaded operating state, and a secondary control instruction is generated and sent to the load control unit.

10. The intelligent operation and maintenance monitoring system for a computer room according to claim 9, characterized in that: The load control unit is used to receive the primary control instructions and the secondary control instructions generated by the degree comparison unit and perform corresponding control operations. The specific contents are as follows: When receiving a first-level control instruction, the execution content is as follows: according to the current server being in an overloaded state, the server cooling fan speed is adjusted; the server load balancing function is automatically triggered, and the backup server equipment and cloud servers are activated to dynamically migrate high-load tasks to the backup server equipment and cloud servers for processing, while restricting the use rights of high-resource-consuming applications; Generate overload operation reports and overload alerts and send them to operation and maintenance personnel; activate the logging function of the central monitoring platform to record the current server operation status information to provide reference for subsequent analysis and optimization; When receiving the secondary control instruction, the execution content is as follows: based on the current server not being in an overloaded operating state, continue to monitor the surface temperature and vibration status of the current server, as well as the server's memory usage and network transmission status; Generate load status reports regularly and store them in the central monitoring platform to provide a reference for subsequent optimization of load assessment models and system adjustments.

Citation Information

Cited By

  • New energy storage remote monitoring system based on cloud computing

    CN120914983A

  • Machine room operation and maintenance fault analysis method using full life cycle monitoring

    CN120928099A

  • Intelligent control method of air conditioning system for temperature control of server room

    CN121665519A