Machine room monitoring method and system based on multi-dimensional data processing
Through the computer room monitoring method based on multi-dimensional data processing, server parameters, line and external environment data are collected and analyzed in real time, and the problems of real-time monitoring and abnormal handling of computer room are solved, achieving rapid and accurate fault location and processing.
Patent Information
- Application Number
- CN202510512843.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
It is difficult for the prior art to realize real-time monitoring of the computer room, and to promptly detect and deal with abnormal nodes, especially in the case of abnormalities of multiple servers for a long time.
The computer room monitoring method based on multi-dimensional data processing is adopted. By acquiring and preprocessing the server's parameter data, line data and external environment parameters, the data before and after the abnormal time nodes are compared with normal data, the possible causes of abnormalities are judged and analyzed, and data is collected in real time when the abnormality occurs for monitoring and analysis.
Real-time monitoring of the computer room is realized, the location and cause of abnormal nodes can be discovered in a timely manner, and the processing is carried out, improving the accuracy and efficiency of troubleshooting.
Smart Images

Figure CN120045376A_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of data processing, and in particular relates to a computer room monitoring method and system based on multi-dimensional data processing. Background Art
[0002] Due to the continuous development of artificial intelligence, the number of intelligent terminal devices has increased rapidly, giving rise to edge computing operations, which in turn has spawned more and more computer rooms composed of edge servers that process data generated by edge terminals. Since the servers in the computer room are mainly used to process real-time data to be processed generated by intelligent terminal devices, this type of data has high latency requirements and requires edge servers to complete data processing in a very short time to obtain the processed results.
[0003] Therefore, in such a computer room where the latency requirements for server data processing are extremely high, real-time monitoring of the computer room is extremely important. If the abnormality cannot be discovered and handled quickly, the abnormal server may not be able to timely process the pending data with extremely high latency requirements generated by the terminal device. In the actual computer room design, the conventional processing method is to perform redundant design, that is, when an abnormality occurs in one of the servers, the task data of the server will be transferred to the redundant server for data processing, so that even if an abnormality occurs in a server, the entire computer room can promptly process the pending data generated by the corresponding smart device terminal within the range.
[0004] The redundant design of the above-mentioned computer room can solve the problem of abnormalities in a short period of time or when a very small number of servers occur, but it cannot solve long-term server abnormalities. In addition, server abnormalities in some cases may also cause abnormalities in other servers. Therefore, how to achieve real-time monitoring of the computer room, timely discover the location and cause of abnormal nodes and deal with them in time is a technical problem that urgently needs to be solved. Summary of the invention
[0005] The purpose of the present invention is to provide a computer room monitoring method and system based on multi-dimensional data processing, so as to realize real-time monitoring of the computer room, timely discover the location and cause of abnormal nodes and handle them in time.
[0006] In order to solve the above technical problems, the technical solution adopted by the present invention is as follows: In a first aspect, a computer room monitoring method based on multidimensional data processing is provided, comprising the following steps: S1: Obtain various parameter data and abnormal data of each server in the computer room, pre-process various data, obtain abnormal time nodes of abnormal data, and extract parameter data before and after the abnormal time nodes from various pre-processed parameter data; S2: Compare the parameter data before and after the abnormal time node with the operation parameter data of other normal operations to determine whether the parameter data before the abnormal time node is abnormal. If so, it is determined as an internal server exception, and perform exception analysis based on the parameter data before and after the abnormal time node. If not, execute step S3; S3: Obtain and preprocess the line data of the server, extract the line data before and after the abnormal time node to determine whether there is a line exception in the server. If so, analyze the line data, the abnormal line position and the abnormal factors. If not, execute step S4; S4: Obtain the external environment parameters of the server, preprocess the external environment parameters of the server, extract the external environment parameters before and after the abnormal time node to determine whether the external environment of the server is abnormal. If so, analyze the external environment exception based on the external environment parameters; S5: Create an internal monitoring module for the servers in the computer room based on the various parameter data of the server, build an external environment monitoring module based on the external environment parameters of the server, and create a line monitoring module based on the line data of the server; S6: When an exception occurs in a certain server, collect the various parameter data of the server in real time and input them into the internal monitoring module. Analyze whether there is an exception in the parameter data of the server through the internal monitoring module. If so, output the analysis result and solution measures of the server parameter exception. If not, collect the line data of the server in real time and transmit it to the line monitoring module to determine whether there is a line exception. If so, output the analysis result of the line exception and give the solution measures. If not, collect the external environment parameters of the server in real time and input them into the external environment monitoring module to analyze whether there is an exception in the external environment parameters of the server. If so, output the analysis result and solution measures of the external environment exception. If not, execute step S7; S7: Inform the management personnel through the warning module to conduct manual fault troubleshooting and fault handling.
[0007] Preferably, the specific process of preprocessing the various data in step S1 is as follows: S11: Construct a filter to remove the noise data in the various data; S12: Obtain the labeled data, and randomly select a specified proportion of the data in the labeled data as the training data to construct a data training subset, and construct an isolation tree with the data training subset; S13: Iteratively execute step S12 to construct a series of isolation trees, divide the denoised various data into multiple data samples, and traverse each data sample in the multiple data samples through each isolation tree in the series of isolation trees; S14: Calculate the average height of each data sample in all the isolation trees, and determine whether each data sample is an outlier according to the average height result.
[0008] Preferably, the external environment parameters of the server include environmental temperature, humidity, and electromagnetic field intensity.
[0009] Preferably, in step S3, obtaining the line data of the server includes voltage data and current data on each line, and the voltage data and current data are respectively collected by a voltage detection module and a current detection module.
[0010] Preferably, the voltage detection module includes a voltage detection circuit, and the voltage detection circuit includes a first resistor, a second resistor, a third resistor, a field effect transistor, a control chip, and a first operational amplifier; The drain of the field effect transistor is connected to the load in the server, the other end of the server load is connected to the power input terminal, the source of the field effect transistor is connected to the third resistor, and the gate of the field effect transistor is connected to the control signal input terminal; One end of the first resistor is connected to the power output terminal, the other end is connected to the second resistor, the IN- port of the control chip is grounded through the second resistor, the VCC port of the control chip is grounded, the IN+ port is grounded through the third resistor, the positive input terminal of the first operational amplifier is connected to the IN- port of the control chip, and the negative input terminal of the first operational amplifier is connected to the IN+ port of the control chip.
[0011] Preferably, the current detection module includes a current detection circuit, and the current detection circuit includes a third resistor, a fourth resistor, a fifth resistor, a sixth resistor, a seventh resistor, an eighth resistor, a ninth resistor, a second operational amplifier, a first capacitor, a second capacitor, a third capacitor, a fourth capacitor, and a fifth capacitor; One end of the third resistor is connected to the voltage input terminal, and the other end is respectively connected to the output terminal of the second operational amplifier, the fourth resistor, and the fifth capacitor. One end of the first capacitor is connected to the voltage input terminal, and the other end is connected to the sixth resistor. The second capacitor is grounded; The positive input terminal of the second operational amplifier is respectively connected to the sixth resistor, the second capacitor, the seventh resistor, and the third capacitor; the negative input terminal of the second operational amplifier is respectively connected to the third capacitor, the fourth capacitor, the eighth resistor, the fourth resistor, and the fifth resistor; After the fifth resistor and the fifth capacitor are connected in series and then connected in parallel with the fourth resistor, the fourth capacitor is connected in parallel across the eighth resistor. After being connected in parallel, one end is connected to the negative input terminal of the second operational amplifier, and the other end is grounded. The seventh resistor is grounded through the ninth resistor.
[0012] In a second aspect, a computer room monitoring system based on multi-dimensional data processing is provided for implementing the computer room monitoring method based on multi-dimensional data processing as described above, including a data acquisition module, a data preprocessing module, a data comparison and analysis module, an internal monitoring module, an external environment monitoring module, a line monitoring module, and an early warning module; The line monitoring module is used to obtain various parameter data, line data, and external environment parameters of each server in the computer room; The data preprocessing module is used to preprocess various data, obtain the abnormal time nodes of the abnormal data, and extract the data before and after the time nodes from the preprocessed parameter data; The data comparison and analysis module is used to compare the data before and after the abnormal time node with other data that has not shown abnormalities, and determine whether there are abnormalities in the parameter data, line data, or external environment data before the abnormal time node; The internal monitoring module is used to analyze whether there are abnormalities in the parameter data of the server. If so, output the analysis results and solutions for the abnormal server parameters; The line monitoring module is used to analyze whether there are abnormalities in the line data of the server. If so, output the analysis results and solutions for the abnormal server line data; The external environment monitoring module is used to analyze whether there are abnormalities in the external environment parameters of the server. If so, output the analysis results and solutions for the abnormal external environment parameters; The warning module is used to inform the management personnel to conduct manual fault troubleshooting and fault handling through the warning module when no abnormalities are detected by the internal monitoring module, external environment monitoring module, and line monitoring module.
[0013] The beneficial effects of the present invention include: The computer room monitoring method and system based on multi-dimensional data processing provided by the present invention establish an automatic fault troubleshooting logic according to historical data, obtain the abnormal time nodes of the abnormal data, extract the parameter data before and after the time nodes, and compare them with the normal operation parameter data to determine whether it is an internal abnormality. Obtain the server line data, extract the line data before and after the abnormal time node to determine whether it is a line abnormality; obtain the external environment parameters of the server to determine whether it is an environmental abnormality; create an internal monitoring module, an external environment monitoring module, and a line monitoring module. When an abnormality occurs in a certain server, conduct troubleshooting based on the fault troubleshooting logic, collect the server parameter data and input it into the internal monitoring module to analyze whether the parameter data is abnormal. Collect the external environment parameters and input them into the external environment monitoring module to analyze whether the external environment parameters are abnormal, and collect the line data and transmit it to the line monitoring module to determine whether there is a line abnormality; if no faults are detected, inform the management personnel to conduct manual fault troubleshooting through the warning module. The above process realizes the real-time monitoring of the computer room. When an abnormality occurs in the server, the location and cause of the abnormal node can be found in time and processed in time. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a schematic flow chart of the computer room monitoring method based on multi-dimensional data processing of the present invention.
[0015] Figure 2 This is a schematic structural diagram of the voltage detection circuit of the present invention.
[0016] Figure 3 This is a schematic structural diagram of the current detection circuit of the present invention.
[0017] Reference numerals: R1 is the first resistor, R2 is the second resistor, R3 is the third resistor, R4 is the fourth resistor, R5 is the fifth resistor, R6 is the sixth resistor, R7 is the seventh resistor, R8 is the eighth resistor, R9 is the ninth resistor, C1 is the first capacitor, C2 is the second capacitor, C3 is the third capacitor, C4 is the fourth capacitor, C5 is the fifth capacitor, P1 is the first operational amplifier, P2 is the second operational amplifier, U is the control chip, VCC is the voltage input terminal, Vout is the voltage output terminal, CTRL is the control signal input terminal, Q is the field effect transistor, D is the drain of the field effect transistor, S is the source of the field effect transistor, and G is the gate of the field effect transistor. Detailed implementation manners
[0018] The following further elaborates on the present invention in conjunction with the attached Figures 1 to 3 drawings: Embodiment 1 Referring to the attached Figure 1 drawings, a computer room monitoring method based on multi-dimensional data processing includes the following steps: S1: Obtain various parameter data and abnormal data of each server in the computer room, preprocess the various data, obtain the abnormal time nodes of the abnormal data, and extract the parameter data before and after the abnormal time nodes from the preprocessed various parameter data; S2: Compare the parameter data before and after the abnormal time node with other operating parameter data that have not shown abnormalities, and determine whether the parameter data before the abnormal time node is abnormal. If so, it is determined as an internal abnormality of the server, and abnormal analysis is performed based on the parameter data before and after the abnormal time node; if not, step S3 is executed; S3: Obtain and preprocess the line data of the server, extract the line data before and after the abnormal time node, and determine whether there is a line abnormality in the server. If so, analyze the line data, the abnormal line position, and the abnormal factors. If not, step S4 is executed; S4: Obtain the external environment parameters of the server, preprocess the external environment parameters of the server, extract the external environment parameters before and after the abnormal time node, and determine whether the external environment of the server is abnormal. If so, analyze the external environment abnormality based on the external environment parameters; S5: Create an internal monitoring module for the servers in the computer room based on the parameter data of the servers, build an external environment monitoring module based on the external environment parameters of the servers, and create a line monitoring module based on the line data of the servers. S6: When an abnormality occurs in a certain server, collect the parameter data of the server in real time and input it into the internal monitoring module. Analyze whether there is an abnormality in the parameter data of the server through the internal monitoring module. If so, output the analysis result of the server parameter abnormality and the solution measures. If not, collect the line data of the server in real time and transmit it to the line monitoring module to determine whether there is a line abnormality. If so, output the analysis result of the line abnormality and give the solution measures; if not, collect the external environment parameters of the server in real time and input them into the external environment monitoring module to analyze whether there is an abnormality in the external environment parameters of the server. If so, output the analysis result of the external environment abnormality and the solution measures. If not, execute step S7. S7: Instruct the management personnel to conduct manual fault troubleshooting and fault handling through the warning module.
[0019] In this embodiment, first, establish an automatic fault troubleshooting logic based on various historical data that may be related to faults, including server parameter data, server line data, and server external environment parameter data. Obtain the abnormal time node of the abnormal data, extract the parameter data before and after this time node, and compare it with the normal operation parameter data to determine whether it is an internal abnormality. Obtain the server line data, extract the line data before and after the abnormal time node to determine a line abnormality; obtain the external environment parameters of the server to determine whether there is an environmental abnormality; create an internal monitoring module, an external environment monitoring module, and a line monitoring module. Then, when an abnormality occurs in a certain server, conduct fault troubleshooting based on the above-created fault troubleshooting logic. First, collect the server parameter data and input it into the internal monitoring module to analyze whether the internal parameter data of the server is abnormal. When no fault point is detected in the internal parameter data, collect the external environment parameters and input them into the external environment monitoring module to analyze whether the external environment parameters are abnormal, and collect the line data and transmit it to the line monitoring module to determine whether there is a line abnormality; if no fault is detected, instruct the management personnel to conduct manual fault troubleshooting through the warning module. The above process realizes the real-time monitoring of the computer room. When an abnormality occurs in the server, it can timely discover the location and cause of the abnormal node and conduct timely processing.
[0020] In another implementation manner of this embodiment, the parameter data such as the CPU temperature, memory usage rate, disk I / O rate, and network traffic of the server are collected in real time by the sensors deployed on the server. At the same time, the voltage and current data of each line are collected by the voltage detection module and the current detection module, and the external environment parameters of the temperature and humidity sensor and the electromagnetic field intensity sensor in the computer room are obtained synchronously. All data is sampled at intervals of 1 second to form a time series data set.
[0021] Smooth the voltage and current data using a preset filter with a cut-off frequency set to 1 kHz to eliminate high-frequency interference. For temperature data, use moving average filtering with a window size of 10 sampling points. Outlier detection is based on the constructed isolation forest model. Randomly select 20% of the historical normal data as the training set and generate 100 isolation trees. For each new data point, calculate the average path depth in all the trees. If the average depth of a data point is lower than the threshold (determined by the 95th percentile of the training set), it is determined as an outlier and removed. Mark the data segments 5 seconds before and after the outlier as candidate areas for abnormal time nodes.
[0022] Taking a certain disk I / O exception as an example, the system detects that the I / O rate suddenly drops to 0. Extract the server parameters (including CPU load, memory occupancy, and process status) from 30 seconds before to 10 seconds after the abnormal time node, and perform dynamic time warping (DTW) comparison with the normal data under the same working conditions in history. It is found that the memory occupancy rate climbs abnormally (exceeding 95% of the threshold) 5 seconds before the exception, and there is no record of new process startup, so it is determined as an internal exception caused by memory leakage. The system automatically triggers memory dump analysis and recommends restarting the relevant services.
[0023] When multiple servers in a certain cabinet report high-temperature alarms simultaneously, collect the environmental data before and after the abnormal time node: The temperature sensor shows that the temperature at the air inlet of the cabinet rises from 22°C to 35°C within 10 seconds, and the humidity drops suddenly from 45% to 30%. The temperatures of other cabinets in the same area are normal; combining with the air-conditioning system log, it is found that the precision air-conditioning compressor in this area fails, resulting in the interruption of the air supply. The system automatically starts the standby air-conditioning unit and adjusts the damper opening of the cold aisle enclosure system.
[0024] For unknown type exceptions (such as intermittent failures caused by motherboard capacitor aging), the warning module notifies the operation and maintenance personnel in the following ways: display a 3D heat map on the monitoring large screen, highlight the location of the abnormal device, push a work order containing a spectrum analysis chart and an abnormal log summary to the mobile terminal, start the AR remote assistance system, and guide the on-site personnel to use a thermal imager to detect specific components.
[0025] Embodiment 2 Based on Embodiment 1, the specific process of preprocessing each item of data in step S1 is as follows: S11: Construct a filter to remove the noise data in each item of data, including irrelevant characters, stop words, and spelling mistakes; extract a part of the data after denoising for data annotation, including extracting specified features from each item of data after denoising, and marking and classifying the data based on the extracted specified features; S12: Obtain the labeled data, randomly select a specified proportion of the data from the labeled data as training data to construct a data training subset, and construct isolation trees with the data training subset; S13: Iteratively execute step S12 to construct a series of isolation trees. Divide the denoised data items into multiple data samples, and traverse each data sample in the multiple data samples through each isolation tree in the series of isolation trees. In each isolation tree, randomly select the specified features and cutting points extracted to split the data. For each leaf node, randomly select a feature and determine a cutting value to split the data into two parts. For each split leaf node, repeat the random split until the maximum number of repeated executions is met, or the isolation tree reaches the preset maximum height, so that only a specified number of data points remain on each leaf node of the isolation tree. Due to the difference between the outliers and the normal points, the outliers will be isolated earlier; S14: Calculate the average height of each data sample in all the isolation trees, and determine whether each data sample is an outlier according to the average height result.
[0026] The above data preprocessing process is the data processing basis of the present invention, providing an accurate data basis for subsequent computer room monitoring. By constructing a series of isolation trees, the denoised data items are divided into multiple data samples, and each data sample in the multiple data samples is traversed through each isolation tree in the series of isolation trees. In each isolation tree, randomly select the specified features and cutting points extracted to split the data. For each leaf node, randomly select a feature and determine a cutting value to split the data into two parts. For each split leaf node, repeat the random split until the maximum number of repeated executions is met, or the isolation tree reaches the preset maximum height, so that only a specified number of data points remain on each leaf node of the isolation tree. Due to the difference between the outliers and the normal points, the outliers will be isolated earlier, realizing the accurate identification of abnormal data. The external environment parameters of the server include environmental temperature, humidity, and electromagnetic field intensity.
[0027] Embodiment 3 Based on Embodiment 1 or Embodiment 2, obtaining the line data of the server in step S3 includes voltage data and current data on each line, and the voltage data and current data are respectively collected by a voltage detection module and a current detection module.
[0028] In this embodiment, the voltage detection module includes a voltage detection circuit, see Figure 2As shown, the voltage detection circuit includes a first resistor R1, a second resistor R2, a third resistor R3, a field effect transistor Q, a control chip U, and a first operational amplifier P1. The drain D of the field effect transistor Q is connected to the load in the server, the other end of the server load is connected to the power input terminal, the source S of the field effect transistor Q is connected to the third resistor R3, and the gate G of the field effect transistor Q is connected to the control signal input terminal CTRL. One end of the first resistor R1 is connected to the power output terminal, and the other end is connected to the first resistor R1. The IN- port of the control chip U is grounded through the first resistor R1, the VCC port of the control chip U is grounded, and the IN+ port is grounded through the third resistor R3. The non-inverting input terminal of the first operational amplifier P1 is connected to the IN- port of the control chip U, and the inverting input terminal of the first operational amplifier P1 is connected to the IN+ port of the control chip U.
[0029] The core component in the voltage detection circuit is the field effect transistor Q. The field effect transistor Q is an NMOS type field effect transistor Q. When performing voltage detection, the characteristics of small impedance when the NMOS transistor is conducting and large impedance when it is not conducting are utilized to achieve voltage detection. When voltage needs to be detected, a high level is output through the input control signal to turn on the NMOS transistor. At this time, the circuit shows a very small resistance, making the ratio of resistor voltage division meet the set ratio, thereby achieving voltage detection.
[0030] The first resistor R1 is connected between VCC and the input terminal of the control chip U. It limits the current to prevent excessive current from damaging the control chip U. At the same time, R1 also plays a role in voltage division, transmitting a part of the input voltage Vout to the input terminal of the control chip U. The second resistor R2 is connected between the output terminal of the control chip U and the ground, forming a feedback network with the output terminal of the operational amplifier P1. The second resistor R2 adjusts the intensity of the feedback signal to help stabilize the working state of the circuit. The third resistor R3 is usually grounded and forms a voltage division circuit with the second resistor R2 to further adjust the voltage level of the feedback signal.
[0031] The field effect transistor Q (N-channel field effect transistor) is used to control the on and off of the current. When the control chip U outputs a high level signal, the voltage of the gate (G) of Q increases, causing Q to conduct, and the current can flow through Q to the load (LOAD). Conversely, when the control chip U outputs a low level signal, Q is cut off and the current is blocked. The control chip U is responsible for detecting the change in the input voltage Vout and controlling the conduction and cutoff of the field effect transistor Q according to the internal logic. It receives the voltage signal from the resistor R1, and judges whether to adjust the output voltage through the internal comparator and logic circuit, thereby controlling the switch state of Q. The first operational amplifier P1 is used for voltage comparison and feedback control. Its inverting input terminal is connected to the ground, and its non-inverting input terminal is connected to the output terminal of the control chip U. When the output voltage of the control chip U changes, P1 will detect this change and adjust the voltage level of the feedback signal through its output terminal, thereby helping to stabilize the working state of the circuit.
[0032] The current detection module includes a current detection circuit, as shown in Figure 3 shown. The current detection circuit includes a third resistor R3, a fourth resistor R4, a fifth resistor R5, a sixth resistor R6, a seventh resistor R7, an eighth resistor R8, a ninth resistor R9, a second operational amplifier P2, a first capacitor C1, a second capacitor C2, a third capacitor C3, a fourth capacitor C4, and a fifth capacitor C5. One end of the third resistor R3 is connected to the voltage input terminal VCC, and the other end is respectively connected to the output terminal of the second operational amplifier P2, the fourth resistor R4, and the fifth capacitor C5. One end of the first capacitor C1 is connected to the voltage input terminal VCC, and the other end is connected to the sixth resistor R6. The second capacitor C2 is grounded; the non-inverting input terminal of the second operational amplifier P2 is respectively connected to the sixth resistor R6, the second capacitor C2, the seventh resistor, and the third capacitor C3; the inverting input terminal of the second operational amplifier P2 is respectively connected to the third capacitor C3, the fourth capacitor C4, the eighth resistor R8, the fourth resistor R4, and the fifth resistor R5. After the fifth resistor R5 and the fifth capacitor C5 are connected in series, they are connected in parallel with the fourth resistor R4. The fourth capacitor C4 is connected in parallel across the eighth resistor R8. After being connected in parallel, one end is connected to the inverting input terminal of the second operational amplifier P2, and the other end is grounded. The seventh resistor is grounded through the ninth resistor R9. Based on the virtual short and virtual open characteristics of the second operational amplifier P2, the circuit converts the current signal into a voltage signal through these characteristics, and then performs current sampling and processing through the analog-to-digital converter.
[0033] The resistor R3 is connected between the input voltage VCC and the non-inverting input terminal of the operational amplifier P2. When current passes through R3, a voltage drop will be generated on it, and this voltage drop is proportional to the current passing through R3. Therefore, R3 plays a role in detecting the input current. The operational amplifier P2 is used as an amplifier in this circuit. Its non-inverting input terminal is grounded through the resistor R4, and its inverting input terminal is connected to the output terminal through the resistor R5 and the feedback network. This configuration enables the operational amplifier to adjust its output voltage according to the change of the input current. The functions of the resistors R4 and R5: The resistors R4 and R5 determine the gain of the operational amplifier.
[0034] Specifically, the gain is the ratio of the resistance values of R4 and R5, taking a negative value because the operational amplifier forms negative feedback here. A change in the input voltage (i.e., the voltage drop across R3) will be amplified into a change in the output voltage. Resistors and capacitors in the feedback network: The sixth resistor R6, the seventh resistor R7, the eighth resistor R8 and the first capacitor C1, the second capacitor C2, the third capacitor C3, the fourth capacitor C4 together constitute the feedback network. This network is not only used to stabilize the output voltage but also plays a filtering role. The sixth resistor R6 and the first capacitor C1, the second capacitor C2 form a high-pass filter for filtering out low-frequency noise; while the seventh resistor R7, the third capacitor C3 and the fourth capacitor C4 form a low-pass filter for filtering out high-frequency noise, making the output signal more stable and less noisy. The ninth resistor R9 is connected between the output terminal of the operational amplifier and the ground to convert the current signal into a voltage signal. When the output voltage of the operational amplifier changes, the current passing through the ninth resistor R9 will also change accordingly, thus generating a voltage drop proportional to the output voltage across the ninth resistor R9. This voltage drop is the voltage signal finally output by the circuit. The first capacitor C1, the second capacitor C2, the third capacitor C3, the fourth capacitor C4 and the fifth capacitor C5 play a filtering role in the circuit. They can eliminate high-frequency noise and interference in the circuit, ensuring the stability and accuracy of the output voltage. The first capacitor C1, the second capacitor C2 and the sixth resistor R6 together form a high-pass filter. The third capacitor C3, the fourth capacitor C4 may form a low-pass filter together with the seventh resistor R7, and the fifth capacitor C5 is used to further stabilize the output signal.
[0035] This current detection circuit realizes the detection and conversion of the input current through various electronic components. The operational amplifier adjusts the output voltage according to the change of the input current and maintains the stability of the output voltage through the feedback network. The capacitors are used for filtering to eliminate noise and ensure the accuracy of the output signal.
[0036] A computer room monitoring system based on multi-dimensional data processing, which is used to implement the described computer room monitoring method based on multi-dimensional data processing, includes a data acquisition module, a data preprocessing module, a data comparison and analysis module, an internal monitoring module, an external environment monitoring module, a line monitoring module, and an early warning module. The line monitoring module is used to obtain various parameter data, line data, and external environment parameters of each server in the computer room. The data preprocessing module is used to preprocess various data, obtain the abnormal time nodes of abnormal data, and extract the data before and after the time nodes from the preprocessed various parameter data. The data comparison and analysis module is used to compare the data before and after the abnormal time node with other data that has not shown abnormalities, and determine whether the parameter data, line data, or external environment data before the abnormal time node is abnormal. The internal monitoring module is used to analyze whether there are abnormalities in the parameter data of the server. If so, it outputs the abnormal analysis result of the server parameters and the solution measures. The line monitoring module is used to analyze whether there are abnormalities in the line data of the server. If so, it outputs the abnormal analysis result of the server line data and the solution measures. The external environment monitoring module is used to analyze whether there are abnormalities in the external environment parameters of the server. If so, it outputs the abnormal analysis result of the external environment parameters and the solution measures. The early warning module is used to inform the management personnel to conduct manual fault troubleshooting and fault handling through the early warning module when no abnormalities are detected through the internal monitoring module, the external environment monitoring module, and the line monitoring module.
[0037] Through multi-dimensional data fusion analysis, the present invention shortens the fault location time to within 30 seconds and improves the accuracy rate to 92%. The specific circuit design achieves a voltage detection accuracy of ±0.5% and a current measurement accuracy of ±1%. Cooperating with a hierarchical abnormal diagnosis process, it effectively distinguishes three types of problems: hardware faults, line abnormalities, and environmental interference, and reduces the false alarm rate by 67%.
[0038] In summary, the computer room monitoring method and system based on multi-dimensional data processing provided by the present invention first establish an automatic fault troubleshooting logic through various historical data and abnormal data, obtain the abnormal time nodes of the abnormal data, extract the parameter data before and after the time nodes, and compare with the normal operation parameter data to determine whether it is an internal abnormality. Obtain the server line data, extract the line data before and after the abnormal time node to determine the line abnormality; obtain the external environment parameters of the server to determine whether the environment is abnormal; create an internal monitoring module, an external environment monitoring module and a line monitoring module. When an abnormality occurs in a certain server, troubleshooting is carried out based on the fault troubleshooting logic, the server parameter data is collected and input into the internal monitoring module to analyze whether the parameter data is abnormal. The external environment parameters are collected and input into the external environment monitoring module to analyze whether the external environment parameters are abnormal, and the line data is collected and transmitted to the line monitoring module to determine whether the line is abnormal; if no fault is detected, the management personnel are informed through the warning module to perform manual fault troubleshooting. The above process realizes the real-time monitoring of the computer room. When an abnormality occurs in the server, the location and cause of the abnormal node can be discovered in time and processed in time.
Claims
1. A computer room monitoring method based on multidimensional data processing, characterized in that: The following steps are involved: S1: Obtain and pre-process various parameter data and abnormal data of the server, obtain the abnormal time node, and extract the parameter data before and after it; S2: Compare the parameter data before and after the abnormal time node with other operating parameter data without abnormality to determine whether the parameter data before the abnormal time node is abnormal. If so, it is considered as an internal abnormality, and abnormality analysis is performed based on the parameter data before and after the abnormal time node; if not, execute step S3; S3: Acquire and pre-process the data of each line, extract the data of each line before and after the abnormal time node to determine whether there is a line abnormality, if so, analyze the abnormal line location and abnormal factors, if not, execute step S4; S4: Obtain external environment parameters, pre-process the external environment parameters, extract the external environment parameters before and after the abnormal time node to determine whether the external environment is abnormal, and if so, analyze the external environment abnormality according to the external environment parameters; S5: Create an internal monitoring module in the computer room based on various parameter data, build an external environment monitoring module based on external environment parameters, and create a line monitoring module based on various line data; S6: When an abnormality occurs, collect various parameter data in real time and input them into the internal monitoring module to analyze whether the parameter data is abnormal. If so, output the abnormality analysis results and solutions. Otherwise, collect various line data in real time and transmit them to the line monitoring module to determine whether there is a line abnormality. If so, output the line abnormality analysis results and provide solutions; otherwise, collect external environment parameters in real time and input them into the external environment monitoring module to analyze whether the external environment parameters are abnormal. If so, output the external environment abnormality analysis results and solutions. If not, execute step S7; S7: Inform management personnel through the early warning module to conduct manual troubleshooting and fault handling.
2. The computer room monitoring method based on multidimensional data processing according to claim 1 is characterized in that: The specific process of preprocessing each data in step S1 is as follows: S11: Construct a filter to remove noise data from each data; S12: Acquire labeled data, and randomly select a specified proportion of data from the labeled data as training data to construct a data training subset, and construct an isolated tree with the data training subset; S13: iteratively executing step S12 to construct a series of isolated trees, dividing each denoised data into multiple data samples, and traversing each data sample in the multiple data samples through each isolated tree in the series of isolated trees; S14: Calculate the average height of each data sample in all isolated trees, and determine whether each data sample is an outlier based on the average height result.
3. The computer room monitoring method based on multidimensional data processing according to claim 1 is characterized in that: The external environmental parameters of the server include ambient temperature, humidity, and electromagnetic field strength.
4. The computer room monitoring method based on multidimensional data processing according to claim 1 is characterized in that: The line data of the server obtained in step S3 includes voltage data and current data on each line, and the voltage data and current data are collected by a voltage detection module and a current detection module respectively.
5. The computer room monitoring method based on multidimensional data processing according to claim 4 is characterized in that: The voltage detection module includes a voltage detection circuit, and the voltage detection circuit includes a first resistor, a second resistor, a third resistor, a field effect transistor, a control chip, and a first operational amplifier; The drain of the field effect tube is connected to the load in the server, the other end of the server load is connected to the power input end, the source of the field effect tube is connected to the third resistor, and the gate of the field effect tube is connected to the control signal input end; One end of the first resistor is connected to the power supply output end, and the other end is connected to the second resistor, the IN- port of the control chip is grounded through the second resistor, the VCC port of the control chip is grounded, and the IN+ port is grounded through the third resistor, the non-inverting input end of the first operational amplifier is connected to the IN- port of the control chip, and the inverting input end of the first operational amplifier is connected to the IN+ port of the control chip.
6. The computer room monitoring method based on multidimensional data processing according to claim 4 is characterized in that: The current detection module includes a current detection circuit, and the current detection circuit includes a third resistor, a fourth resistor, a fifth resistor, a sixth resistor, a seventh resistor, an eighth resistor, a ninth resistor, a second operational amplifier, a first capacitor, a second capacitor, a third capacitor, a fourth capacitor, and a fifth capacitor; One end of the third resistor is connected to the voltage input end, and the other end is respectively connected to the output end of the second operational amplifier, the fourth resistor and the fifth capacitor; one end of the first capacitor is connected to the voltage input end, and the other end is connected to the sixth resistor; the second capacitor is grounded; The non-inverting input terminal of the second operational amplifier is respectively connected to the sixth resistor, the second capacitor, the seventh resistor and the third capacitor; the inverting input terminal of the second operational amplifier is respectively connected to the third capacitor, the fourth capacitor, the eighth resistor, the fourth resistor and the fifth resistor; The fifth resistor and the fifth capacitor are connected in series and then connected in parallel with the fourth resistor. The fourth capacitor is connected in parallel to both ends of the eighth resistor. After the parallel connection, one end is connected to the inverting input end of the second operational amplifier and the other end is grounded. The seventh resistor is grounded through the ninth resistor.
7. A computer room monitoring system based on multidimensional data processing, used to implement a computer room monitoring method based on multidimensional data processing according to any one of claims 1 to 6, characterized in that: It includes data acquisition module, data preprocessing module, data comparison and analysis module, internal monitoring module, external environment monitoring module, line monitoring module and early warning module; The line monitoring module is used to obtain various parameter data, line data and external environment parameters of each server in the computer room; The data preprocessing module is used to preprocess various data, obtain abnormal time nodes of abnormal data, and extract various data before and after the time nodes from various parameter data after preprocessing; The data comparison and analysis module is used to compare the data before and after the abnormal time node with other data without abnormality, and determine whether the parameter data, line data or external environment data before the abnormal time node is abnormal; The internal monitoring module is used to analyze whether the parameter data of the server is abnormal, and if so, output the server parameter abnormality analysis results and solutions; The line monitoring module is used to analyze whether there is an abnormality in the line data of the server, and if so, output the abnormality analysis result of the server line data and the solution; The external environment monitoring module is used to analyze whether there are any abnormalities in the external environment parameters of the server, and if so, output the abnormal analysis results of the external environment parameters and the solution measures; The early warning module is used to inform the management personnel to perform manual troubleshooting and fault handling through the early warning module when no abnormality is found through the internal monitoring module, the external environment monitoring module, and the line monitoring module.
Citation Information
Patent Citations
Equipment monitoring system and method
CN116800806A
Machine room operation state detection method and device, storage medium and electronic equipment
CN117891234A
Multi-dimensional anomaly detection method and device based on fusion model
CN118094383A
Machine room information data management method and system
CN119179704A
Big data-based computer room inspection method and related device
WO2020143327A1
Cited By
Machine room environment multi-dimensional intelligent monitoring system and method based on edge calculation
CN120256252A
Multi-dimensional intelligent monitoring system and method for computer room environment based on edge computing
CN120256252B