A server state intelligent early warning method and system based on a harmonic correlation model
By employing a harmonic correlation model in data centers, a smart early warning method is used to achieve deep coupling analysis of power quality and server status. This accurately pinpoints the causes of performance degradation or overheating due to harmonics, reduces blind troubleshooting during operation and maintenance, and improves operational efficiency and data center stability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CHANGFENG INNOVATION TECHNOLOGY CO LTD
- Filing Date
- 2026-03-23
- Publication Date
- 2026-06-26
AI Technical Summary
In existing technologies, the power quality and server status monitoring of data centers lack effective coordination and correlation analysis, making it difficult for maintenance personnel to quickly determine the root cause of harmonic problems causing server performance degradation or abnormal temperature rise. Static threshold alarms cannot reflect the complex nonlinear relationship between harmonics and server status, which may miss hidden performance degradation risks or generate false alarms.
A server status intelligent early warning method based on a harmonic correlation model is adopted. By synchronously collecting power supply harmonic data and server status data, a multi-source dataset with timestamps is generated. The dataset is then aligned with timestamps, preprocessed, and subjected to machine learning-driven correlation analysis to generate a harmonic impact baseline curve. The server status is monitored in real time, and the baseline curve is used to determine the cause of the harmonics. A graded early warning report is generated, and the parameters of the correlation analysis model are optimized.
It enables deep coupling analysis of harmonics and server status, accurately locates potential root causes, reduces the time spent on blind troubleshooting in operation and maintenance, prevents hidden performance degradation and sudden downtime risks, improves the accuracy and initiative of data center operation and maintenance, and ensures the long-term effectiveness and adaptability of the system.
Smart Images

Figure CN122285442A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data center infrastructure management technology, and in particular to a method and system for intelligent early warning of server status based on a harmonic correlation model. Background Technology
[0002] With the rapid development of cloud computing, big data, and artificial intelligence, the scale and server density of data centers continue to grow, and the complexity and reliability requirements of their operation have reached unprecedented levels. In data center infrastructure, power supply and cooling systems are the cornerstones for ensuring the stable operation of IT equipment. Power quality, especially power harmonic pollution introduced by nonlinear loads, has become a significant potential threat. Harmonics can cause voltage waveform distortion, increase line losses, cause equipment overheating, and may interfere with the normal operation of precision electronic equipment. Meanwhile, the performance and temperature status of the server's core processor (CPU) are direct indicators of its health and computing efficiency. Currently, the industry typically uses independent monitoring systems to separately monitor the power quality (including harmonics) and the performance and temperature of IT equipment in the data center, forming separate data streams and alarm systems.
[0003] However, in practical applications, due to the lack of effective coordination and correlation analysis between power supply harmonic monitoring and server status monitoring at the data level, when servers experience performance degradation or abnormal temperature increases, maintenance personnel find it difficult to quickly determine whether the root cause is a harmonic problem in the power supply system. Existing early warning mechanisms typically rely on setting static thresholds for single parameters, such as triggering an alarm when the CPU temperature exceeds a certain fixed value or the total harmonic distortion rate exceeds a certain standard limit. However, this static threshold alarm method cannot reflect the complex and non-linear dynamic correlation between harmonics and server status. It may miss the hidden performance degradation risk caused by slow harmonic deterioration, or generate false alarms when harmonics fluctuate instantaneously but do not have a substantial impact on the server.
[0004] Therefore, how to achieve a deep correlation analysis between power supply quality and server operating status, and to perform intelligent root cause determination and graded response when anomalies occur, thereby improving the accuracy and proactivity of data center operation and maintenance, is a technical problem that urgently needs to be solved. Summary of the Invention
[0005] To address the aforementioned technical issues, this application provides a server status intelligent early warning method and system based on a harmonic correlation model.
[0006] Firstly, this application provides a server status intelligent early warning method based on a harmonic correlation model, employing the following technical solution: Synchronously collect power harmonic data and server status data from the data center to generate a raw multi-source dataset with timestamps; the server status data includes server CPU performance data and server temperature data. The original multi-source dataset is timestamped and preprocessed to generate a standardized analysis dataset. The correlation analysis model is invoked to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power harmonic data and server status data, and generate a harmonic impact baseline curve. The server status data is monitored in real time. When the server status data exceeds the preset warning threshold, the harmonic influence reference curve is called to determine the cause and generate an anomaly determination report. Based on the anomaly determination report, a tiered early warning report is generated and pushed to the designated terminal; Based on the tiered early warning report, operation and maintenance data are recorded, and the parameters of the correlation analysis model are optimized to update the harmonic impact baseline curve.
[0007] By adopting the above technical solution, traditionally fragmented power quality monitoring and IT equipment health monitoring are deeply coupled. Machine learning is used to reveal their inherent correlations, enabling the system to not only issue alerts when servers experience performance degradation or overheating, but also accurately pinpoint the potential root cause—power harmonics—and provide differentiated early warnings and remediation guidance accordingly. This technical solution significantly reduces the time spent by maintenance personnel on blind troubleshooting, preventing hidden performance degradation and sudden downtime risks caused by harmonic issues at the source, and improving the overall operational efficiency and business continuity of the data center. Simultaneously, the system's self-learning optimization capabilities ensure its long-term effectiveness and adaptability, realizing a transformation in data center operations from passive response to proactive prediction, and from experience-driven to data intelligence-driven.
[0008] Optionally, the steps of timestamp alignment and preprocessing of the original multi-source dataset to generate a standardized analysis dataset include: Based on the timestamps of the original multi-source dataset, power harmonic data, server CPU performance data, and server temperature data are aligned and integrated to generate an aligned and integrated dataset. The aligned and integrated dataset is cleaned to remove outliers and generate a cleaned dataset. The cleaned dataset is formatted and the data units are standardized to a preset standard to generate a standardized dataset. The standardized dataset is normalized to eliminate the differences in the numerical magnitude of different parameters, thereby generating a standardized analysis dataset.
[0009] By adopting the above technical solutions, the original multi-source data is aligned with timestamps to ensure the temporal accuracy of causal analysis, data cleaning ensures the authenticity and reliability of data samples, format standardization achieves unambiguous interpretation and computability of data, and finally, normalization provides fair and unbiased input features for machine learning models, ensuring that the analytical conclusions relied upon by the entire intelligent early warning system are based on high-quality and highly consistent data, thereby improving the accuracy and credibility of early warning.
[0010] Optionally, the steps of calling the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculating the correlation weights between power supply harmonic data and server status data, and generating a harmonic impact baseline curve include: Call the pre-trained correlation analysis model and input the power harmonic data and server status data from the standardized analysis dataset; The first correlation weight between power supply harmonic data and server CPU performance data, and the second correlation weight between power supply harmonic data and server temperature data are calculated using machine learning algorithms to generate a correlation weight matrix. Based on the aforementioned correlation weight matrix, a mapping relationship is constructed between the power supply harmonic content range and the server CPU performance attenuation range, as well as a mapping relationship between the power supply harmonic content range and the server temperature rise range, generating a harmonic influence baseline curve.
[0011] By adopting the above technical solution, a pre-trained model is invoked to achieve rapid startup. By calculating the correlation weight matrix, the complex nonlinear relationship between harmonics and server status is quantified in depth, thereby constructing an intuitive and usable harmonic impact benchmark curve, transforming the black-box output of machine learning into a white-box decision-making tool for operation and maintenance.
[0012] Optionally, the steps of monitoring the server status data in real time, and when the server status data exceeds a preset warning threshold, calling the harmonic influence reference curve to determine the cause and generating an anomaly determination report include: Real-time monitoring of server status data; when server CPU performance data or server temperature data exceeds the preset warning threshold, obtain the power harmonic data at the current moment. Call the harmonic impact baseline curve to extract the server status data impact range corresponding to the current power harmonic data; Compare the current server status data change range with the range of the influence interval, and analyze whether the server status data change trend is consistent with the predicted trend of the harmonic influence benchmark curve. If the change in the current server status data is within the influence range and the trend of the server status data change is consistent with the predicted trend of the harmonic influence baseline curve, then it is determined to be a harmonic-dominated influence; otherwise, it is determined to be influenced by other factors. Generate an anomaly assessment report, including the assessment results, current power harmonic data, and the magnitude of changes in server status data.
[0013] By adopting the above technical solution, efficient triggering is achieved based on preset warning thresholds. Real-time data is input into a historical knowledge model by calling the harmonic impact benchmark curve. Rigorous hypothesis testing is conducted through dual comparison of amplitude and trend, and finally, a judgment result is output based on clear rules, generating a comprehensive anomaly judgment report. This technical solution not only significantly improves fault location efficiency and reduces blind troubleshooting by maintenance personnel, but more importantly, by effectively distinguishing between anomalies caused by harmonics and those not caused by harmonics, it can precisely guide limited maintenance resources towards the correct handling direction. This fundamentally prevents server performance degradation and downtime risks caused by power quality issues, ensuring the ultra-high reliability and stability of data center operations.
[0014] Optionally, the step of generating a tiered early warning report and pushing it to a designated terminal based on the anomaly determination report includes: Analyze the anomaly determination report to extract the determination result, current power harmonic data, and server status data change amplitude; Based on the judgment result and the magnitude of changes in server status data, a preset graded early warning rule is matched to determine the early warning level; Based on the warning level, a graded warning report containing handling suggestions is generated and sent to the designated terminal through a push interface.
[0015] By adopting the above technical solutions, a highly automated, strategy-driven, and precisely targeted early warning information generation and distribution system is constructed. It achieves intelligent and standardized response strategies through tiered early warning rule matching, provides action guidelines containing actionable knowledge that match the severity of the event through tiered report generation, and finally ensures that instructions are accurately and promptly delivered to the responsible parties through interface call push.
[0016] Optionally, the steps of recording operation and maintenance data based on the tiered early warning report, optimizing the parameters of the correlation analysis model, and updating the harmonic impact baseline curve include: Receive tiered early warning reports, extract the early warning level, the magnitude of changes in server status data, and handling suggestions; Based on the proposed handling suggestions, obtain the handling operation data performed by the operation and maintenance personnel and the change range of server status data after the handling, generate an operation and maintenance handling record dataset and synchronize it to the historical database; Call the correlation analysis model, load the operation and maintenance record dataset from the historical database, and optimize the correlation analysis model parameters; Based on the optimized correlation analysis model, the correlation weight matrix and harmonic influence baseline curve are updated.
[0017] By adopting the above technical solution, the system records handling operations and results to digitize human experience. This data is then synchronized to a historical database for continuous knowledge accumulation. Finally, by calling the model and loading new data to optimize parameters, the system can learn from each operational practice, verify hypotheses, and correct its understanding. This closed loop ensures that the system's core analysis model does not stagnate but evolves continuously with data center equipment updates, load changes, and the implementation of governance measures, resulting in increasingly accurate harmonic impact baseline curves. This technical solution endows the solution with long-term vitality and scenario adaptability, fundamentally guaranteeing that early warning accuracy continuously improves over time. It provides the technical support for enabling data center operations to move from "experience-driven" to "data intelligence-driven" and possess "continuous learning" capabilities.
[0018] Optionally, the early warning method further includes: Real-time collection of server load data, generating a server load dataset with timestamps; The server load dataset is timestamped and aligned with the standardized analysis dataset to generate a time-synchronized dataset. Perform data integration processing on the time-synchronized dataset to generate an extended analysis dataset; The association analysis model is invoked to perform machine learning-driven association analysis on the extended analysis dataset, calculate the association weights of server load parameters, power harmonic data and server status data, and generate a comprehensive influence baseline curve. Based on the comprehensive impact benchmark curve, the preset early warning threshold is dynamically adjusted; The server status data is monitored in real time, and the changing trend characteristics of the current server load data are extracted. When the server status data exceeds the adjusted preset warning threshold, the cause is determined by combining the comprehensive impact benchmark curve and the trend characteristics, and an optimized anomaly determination report is generated.
[0019] By adopting the above technical solution, the system can clearly distinguish whether server anomalies stem from inherent changes in its internal workload, external power quality (harmonic) interference, or a combination of both, thus resolving the ambiguity in the "other factors" determination in the original solution. By dynamically adjusting the warning threshold, the system intelligently adapts to the natural fluctuations in server load, reducing false alarm rates and enhancing the sensitivity to detect hidden problems. Combined with the analysis of load change trend characteristics, the system can not only diagnose existing anomalies but also perceive emerging risk patterns. The resulting optimized anomaly report provides unprecedented root cause transparency and action guidance.
[0020] Secondly, this application provides a server status intelligent early warning system based on a harmonic correlation model, which adopts the following technical solution: The multi-source data synchronous acquisition module is used to synchronously acquire power harmonic data and server status data from the data center, generating a raw multi-source dataset with timestamps; among which, the server status data includes server CPU performance data and server temperature data. The data processing module is used to perform timestamp alignment and preprocessing on the original multi-source datasets to generate standardized analysis datasets; The baseline curve generation module is used to call the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power harmonic data and server status data, and generate a harmonic influence baseline curve. The anomaly detection module is used to monitor the server status data in real time. When the server status data exceeds the preset warning threshold, the module calls the harmonic influence reference curve to determine the cause and generates an anomaly detection report. The graded early warning module is used to generate a graded early warning report based on the anomaly determination report and push it to the designated terminal; The feedback optimization module is used to record operation and maintenance data based on the hierarchical early warning report and optimize the parameters of the correlation analysis model.
[0021] Thirdly, this application provides a computer device, which adopts the following technical solution: A computer device includes a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to perform the steps of the method as described in the first aspect.
[0022] Fourthly, this application provides a computer-readable storage medium, which adopts the following technical solution: A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the methods in the first aspect. Attached Figure Description
[0023] Figure 1 This is a first flowchart illustrating a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application.
[0024] Figure 2 This is a second flowchart illustrating a server status intelligent early warning method based on a harmonic correlation model, according to one embodiment of this application.
[0025] Figure 3 This is a schematic diagram of the third process of a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application.
[0026] Figure 4 This is a schematic diagram of the fourth process of a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application.
[0027] Figure 5 This is a schematic diagram of the fifth process of a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application.
[0028] Figure 6 This is a schematic diagram of the sixth process of a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application.
[0029] Figure 7 This is a schematic diagram of the seventh process of a server status intelligent early warning method based on a harmonic correlation model, which is one embodiment of this application. Detailed Implementation
[0030] To make the purpose, technical solution, and advantages of this application clearer, the following description is provided in conjunction with the appendix. Figures 1-7 The present application will be further described in detail below with reference to embodiments. It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the application.
[0031] This application discloses an intelligent early warning method for server status based on a harmonic correlation model.
[0032] Reference Figure 1 A server status intelligent early warning method based on a harmonic correlation model, specifically including: Step S101: Synchronously collect power harmonic data and server status data from the data center to generate a raw multi-source dataset with timestamps; wherein, the server status data includes server CPU performance data and server temperature data. Specifically, in traditional data center monitoring, power quality (harmonics), computing core performance (CPU), and hardware health status (temperature) are usually monitored by independent systems. The data is fragmented in time and space, making it impossible to conduct effective causal correlation analysis.
[0033] In this embodiment, by deploying a harmonic analyzer, using a server out-of-band management interface (such as IPMI / BMC), and built-in / external temperature sensors, and using a unified acquisition clock (usually synchronized by a central server or network time protocol) as a reference, these three types of heterogeneous data are forced to be assigned the same time identifier upon generation. Power supply harmonic data (such as total harmonic distortion (THD), harmonic currents and voltages) reflects the purity of energy supply; CPU performance data (such as utilization, processing speed, and instruction latency) reflects the execution efficiency of computing tasks; and server temperature data (such as CPU core temperature and chassis temperature) reflects the physical state of hardware operation.
[0034] Next, a unified timestamp is bound to these three, which is equivalent to establishing a traceable and strictly corresponding spatiotemporal coordinate for each power supply disturbance, each performance fluctuation, and each temperature rise change. This integrates the originally isolated parameter streams into a raw multi-source dataset that can reflect "how power supply quality affects the computing unit in real time and is ultimately manifested as physical changes," providing the possibility for subsequent deep correlation analysis.
[0035] Step S102: Perform timestamp alignment and preprocessing on the original multi-source dataset to generate a standardized analysis dataset; Although the data is timestamped during acquisition, there may be slight differences in acquisition cycles and communication delays between different devices. Timestamp alignment is achieved through interpolation, resampling, and other techniques to ensure that at any given point in the analysis, the three sets of data—harmonics, performance, and temperature—strictly correspond to the same physical moment. This is a prerequisite for accurate correlation analysis.
[0036] Subsequently, the preprocessing steps include data cleaning, format standardization, and normalization. Data cleaning aims to remove outliers caused by momentary sensor malfunctions, communication packet loss, or strong electromagnetic interference (such as sudden temperature jumps to unreasonable values), ensuring data authenticity. Format standardization converts data from different manufacturers and protocols (e.g., temperature units may be Celsius or Fahrenheit, and current units may be amperes or milliamperes) into a unified internal measurement standard, eliminating ambiguity. Normalization maps parameters with vastly different orders of magnitude (e.g., THD values may be single-digit percentages, while CPU temperatures are tens or hundreds of degrees Celsius) to similar numerical ranges (e.g., [0,1]) through mathematical transformations. This prevents features with large values from "overwhelming" features with small values during subsequent machine learning model training, ensuring the model can fairly learn the influence of all parameters.
[0037] After this series of processing steps, the resulting standardized analysis dataset is a clean, consistent, high-quality data set that can be directly used for complex mathematical operations.
[0038] Step S103: Call the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power supply harmonic data and server status data, and generate a harmonic impact baseline curve. Among these features, the system autonomously learns and quantifies the hidden patterns of the impact of power supply harmonics on server status from historical data. Simple threshold monitoring or empirical formulas cannot characterize the nonlinear and coupled relationship that may exist between harmonics (a complex spectrum) and multidimensional server status (performance, temperature).
[0039] Therefore, this embodiment employs machine learning algorithms (such as random forests and gradient boosting trees), which can handle high-dimensional data and automatically find patterns from a "standardized analysis dataset." Specifically, the model is trained using harmonic parameters (such as THD, 3rd and 5th harmonic content) as input features and server status data (such as CPU processing speed decrease and core temperature increase) as prediction targets. During training, the algorithm iterates repeatedly, calculating correlation weights, which essentially quantifies the degree of influence of each harmonic component on each server status indicator. For example, the model might find that the 5th harmonic current has the highest weight on CPU instruction latency, while the 7th harmonic has a more significant impact on chassis temperature rise.
[0040] Ultimately, the patterns learned by the model are solidified into one or more harmonic influence baseline curves. This curve is not a simple formula, but a complex mapping relationship containing probability distributions. It can intuitively or through querying show how likely it is that when the total harmonic distortion rate reaches a certain range, it will cause the CPU processing speed to decrease by a certain percentage or the core temperature to rise by a certain degree Celsius.
[0041] Step S104: Monitor server status data in real time. When the server status data exceeds the preset warning threshold, call the harmonic influence reference curve to determine the cause and generate an anomaly determination report. This implementation utilizes an established quantitative knowledge model to perform intelligent attribution analysis on real-time faults. Traditional monitoring can only issue an alarm stating "Server temperature exceeds limit" when CPU temperature is too high or performance degrades, requiring maintenance personnel to rely on experience to troubleshoot various possible causes such as heat dissipation, load, and power supply. This embodiment, however, presets warning thresholds based on server design specifications and business requirements (such as CPU core temperature of 85°C or a decrease in processing speed exceeding 10%).
[0042] When the real-time monitored server status data reaches these thresholds, the system does not immediately conclude that it is a heat dissipation problem. Instead, it initiates an intelligent diagnostic process: it immediately retrieves the power harmonic data at the current moment (after timestamp alignment) and compares it with the harmonic impact baseline curve. The logic for determining the cause is as follows: if the current harmonic content falls exactly within the range predicted by the baseline curve to lead to the currently observed performance degradation or temperature rise, and the trend of parameter changes (e.g., a slow increase in harmonics accompanied by a synchronous slow increase in temperature) is consistent with historical patterns, then the system has a high degree of confidence in determining that the anomaly is mainly caused by harmonics. Conversely, if the harmonic content is very low and within the "safe zone," it is determined to be due to other factors (such as fan failure, cold aisle blockage, sudden high load, etc.). The final anomaly report not only includes the abnormal phenomenon but also the diagnostic conclusion of the cause of the anomaly, which fundamentally changes the information value of the early warning.
[0043] Step S105: Based on the anomaly determination report, generate a graded early warning report and push it to the designated terminal; The logic behind this step is to implement differentiated response strategies based on the root cause and severity of the fault, thereby optimizing the scheduling of operation and maintenance resources.
[0044] Specifically, based on the "cause" (whether it is caused by harmonics) and "severity" (how much the parameters exceed the standard) clearly stated in the anomaly assessment report, the system activates a tiered early warning mechanism. For example, a Level 1 warning is for minor harmonic exceedances and minor server anomalies; the report may only be pushed to frontline inspection personnel to alert them. A Level 2 warning is for moderate harmonic exceedances that have caused perceptible performance degradation; the report will be pushed to the operations and maintenance supervisor along with preliminary remediation suggestions (such as "it is recommended to check for harmonic sources in the A-column cabinet or install filters"). A Level 3 warning is for severe harmonic exceedances that pose a risk of server downtime; the report will be pushed to the operations and maintenance team and even management through multiple channels such as SMS, email, and monitoring dashboards, requiring immediate intervention.
[0045] Understandably, this tiered early warning mechanism ensures that the seriousness of the alert matches the urgency of the response, avoiding alert fatigue and ensuring that critical risks are addressed with the highest priority. Pushing alerts to designated terminals signifies precise information routing, allowing personnel in different roles to receive the information most needed within their respective areas of responsibility.
[0046] Step S106: Based on the graded early warning report, record the operation and maintenance handling data, optimize the parameters of the correlation analysis model, and update the harmonic impact baseline curve.
[0047] When maintenance personnel take actual actions based on the graded early warning report (such as installing active filters or adjusting load distribution), the system requires or allows the recording of the handling process (what equipment was used and which line was treated) and the handling results (harmonic values after treatment, server performance and temperature recovery) as maintenance handling data. These data are extremely valuable "treatment-feedback" cases.
[0048] Next, the system feeds this new data (data from the entire process of anomaly generation to governance and recovery) back into the correlation analysis model. Using this new real-world feedback data, the model retrains or fine-tunes its parameters, thereby updating the harmonic impact baseline curve. For example, it might discover that a new type of server is more sensitive to a certain harmonic than historical models predicted, and the model will then make corrections. Through this continuous optimization cycle, the system's diagnostic model evolves with changes in data center equipment and load patterns, becoming increasingly accurate, forming a virtuous cycle of becoming smarter with use.
[0049] The above implementation deeply couples traditionally fragmented power quality monitoring with IT equipment health monitoring, leveraging machine learning to reveal their inherent correlations. This allows for not only alerts when server performance degrades or overheating occurs, but also precise identification of power harmonics as a potential root cause, providing differentiated early warnings and remediation guidance. This technical solution significantly reduces the time spent by maintenance personnel on blind troubleshooting, preventing hidden performance degradation and sudden downtime risks caused by harmonic issues at the source, and improving the overall operational efficiency and business continuity of the data center. Simultaneously, the system's self-learning optimization capabilities ensure its long-term effectiveness and adaptability, realizing a paradigm shift in data center operations from passive response to proactive prediction, and from experience-driven to data intelligence-driven.
[0050] Reference Figure 2 As one implementation of step S102, the steps of timestamp alignment and preprocessing of the original multi-source dataset to generate a standardized analysis dataset include: Step S201: Based on the timestamps of the original multi-source dataset, align and integrate the power harmonic data, server CPU performance data, and server temperature data to generate an aligned and integrated dataset. Although harmonic, performance, and temperature data are all timestamped at the acquisition end, slight clock drifts may exist within different acquisition devices (such as independent harmonic analyzers, server BMC chips, and external temperature sensors), or random delays may occur during data transmission to the processing unit via the network. This results in millisecond or even second-level discrepancies in the timestamps recorded for the three types of data, theoretically corresponding to the same physical moment. Directly using this "asynchronous" data for analysis will lead to serious misjudgments of causality; for example, incorrectly associating harmonic distortion from a previous moment with a CPU temperature spike in a subsequent moment.
[0051] Therefore, the alignment and integration in this embodiment is not a simple accumulation of time points, but rather a forced alignment of these three data streams onto a unified time axis through interpolation, resampling, or matching algorithms based on the nearest neighbor timestamp. This ensures that within any analyzed time slice, the harmonic data, CPU performance data, and temperature data strictly correspond to the actual state at the same physical instant. The generated aligned and integrated dataset is a prerequisite for all subsequent analyses; it is equivalent to constructing a fully synchronized, multi-dimensional digital twin state snapshot sequence for the data center in the time dimension.
[0052] Step S202: Clean the aligned and integrated dataset, remove outliers, and generate a cleaned dataset. Even when data is time-aligned, it may still contain invalid or distorted information. Outliers typically arise from various situations: for example, a temperature sensor outputting a -40°C jump due to momentary poor contact; a harmonic analyzer capturing a very brief spike in power supply quality during an electrical switch, which does not represent steady-state power quality; or data packet errors in network communication causing severe distortion of performance data. These points are not the true system state, but rather "false signals" during measurement or transmission. If these outliers are retained, they will be treated as valid samples in subsequent statistical analysis and machine learning model training, severely distorting the data distribution and causing the model to learn incorrect patterns (e.g., mistaking a low temperature point caused by a sensor malfunction as evidence of effective harmonic mitigation).
[0053] In this embodiment, the data cleaning process automatically identifies and removes outliers by setting threshold ranges based on physical common sense (e.g., the server CPU core temperature cannot be lower than room temperature or higher than the semiconductor junction temperature limit), outlier detection based on statistical laws (e.g., the 3-sigma principle), or smoothing filtering algorithms based on data continuity. The resulting cleaned dataset eliminates most of the data noise, making the dataset more reflective of the true operating status of the monitored system and laying a clean data foundation for subsequent accurate analysis.
[0054] Step S203: Standardize the format of the cleaned dataset, unify the data units to the preset standard, and generate a standardized dataset; The raw data from different manufacturers and models of equipment often have different numerical representation formats and units. For example, temperature data might come from a sensor from manufacturer A, expressed in degrees Celsius (°C) as a floating-point number; while data from a sensor from manufacturer B might be expressed in one-tenth of a degree Fahrenheit as an integer. In harmonic data, total harmonic distortion (THD) might be expressed as a percentage (%) or in per-unit values (pu); current values might be in amperes (A) or milliamperes (mA). If this data with inconsistent units and mixed formats is directly input into the mathematical model, the calculations will completely lose their physical meaning and may even lead to program errors.
[0055] In the embodiments of this application, format standardization is to perform a mandatory conversion process. Based on a preset standard (usually a standard defined internally by the system, such as temperature being standardized to degrees Celsius, harmonics to percentages, and current to amperes), all data items in the "cleaned dataset" are converted to a consistent unit of measurement and data structure through a determined conversion formula and data type conversion.
[0056] The standardized dataset generated in this step ensures that each value has a clear and consistent meaning, enabling direct comparison, addition, subtraction, weighting, and other mathematical operations on data from different sources. This is a prerequisite for any quantitative analysis.
[0057] Step S204: Normalize the standardized dataset to eliminate the differences in the numerical magnitude of different parameters and generate a standardized analysis dataset.
[0058] In standardized datasets, although units are unified, the numerical ranges (orders of magnitude) of different physical parameters can vary greatly. For example, CPU core temperature typically ranges from tens to hundreds (e.g., 30-90°C), while total harmonic distortion (THD) might only be a single-digit percentage (e.g., 1.5%-8.0%), and CPU utilization is a percentage between 0-100%. If these raw features with vastly different orders of magnitude are directly input into machine learning algorithms (such as random forests, gradient boosting trees, etc.), features with larger numerical ranges (e.g., temperature) will naturally dominate in calculating distances and determining split points. Small absolute changes in temperature may overshadow significant changes in features with smaller numerical ranges (e.g., THD), causing the model to fail to accurately learn the latter's impact.
[0059] In this embodiment, normalization can be achieved through mathematical transformations (such as min-max normalization, which linearly scales the data to the [0,1] interval; or Z-score standardization, which converts the data into a distribution with a mean of 0 and a standard deviation of 1), mapping all features to the same approximately identical numerical scale. This process eliminates the differences in the original numerical magnitudes while preserving the distributional relationships and relative sizes within the data.
[0060] Ultimately, the generated standardized analysis dataset has each feature dimension in a similar numerical range. This allows the subsequent association analysis model to be unaffected by the dimensions and fairly evaluate the true contribution of each input feature (whether it is temperature, harmonics, or performance indicators) to the output target (such as performance degradation), thereby learning more accurate and robust association patterns.
[0061] In the above implementation, the original multi-source data is timestamped to ensure the temporal accuracy of causal analysis, data cleaning ensures the authenticity and reliability of data samples, format standardization achieves unambiguous interpretation and computability of data, and normalization provides fair and unbiased input features for machine learning models. This ensures that the analytical conclusions relied upon by the entire intelligent early warning system are based on high-quality and highly consistent data, thereby improving the accuracy and credibility of early warning.
[0062] Reference Figure 3 As one implementation of step S103, the steps of calling the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculating the correlation weight between power supply harmonic data and server status data, and generating a harmonic impact baseline curve include: Step S301: Call the pre-trained correlation analysis model and input the power harmonic data and server status data from the standardized analysis dataset; The pre-trained association analysis model is not a blank or randomly initialized model, but a machine learning model that has been preliminarily trained using large-scale historical operational data (e.g., harmonic, performance, and temperature data of the data center over the past months or even years). This pre-training allows the model to form preliminary prior knowledge about the potential relationship patterns between these parameters, enabling it to adjust and infer on new data faster and more stably compared to training from scratch.
[0063] Next, by inputting parameters from the standardized analysis dataset into the model, we are essentially submitting high-quality evidence that has undergone rigorous cleaning, alignment, and scaling to this "data analysis expert." Power supply harmonic data (such as the content of each harmonic and total harmonic distortion (THD)) are input as potential "causal variables" (features), while server status data (CPU performance parameters such as instruction latency and processing speed, and temperature parameters such as core temperature) are input as "outcome variables" (labels). The aim is to request the model to quantify the impact of the former on the latter. This step is the starting point for all subsequent quantitative correlation analyses.
[0064] Step S302: Calculate the first correlation weight between power supply harmonic data and server CPU performance data, and the second correlation weight between power supply harmonic data and server temperature data using a machine learning algorithm, and generate a correlation weight matrix. The logic behind this step lies in deconstructing and quantifying the nonlinear influence strength between multiple causes and effects in a complex system. The impact of power supply harmonics on servers is complex and coupled: different harmonic orders (such as the 3rd, 5th, and 7th) affect CPU computing efficiency (performance) and chip heat generation (temperature) in different ways and to varying degrees. Simple correlation analysis cannot capture this multivariate, nonlinear relationship.
[0065] In this embodiment, machine learning algorithms (such as random forests and gradient boosting trees) play a role. They learn patterns from the input data through complex internal operations (such as building a large number of decision trees and integrating their judgments). One of its core outputs is the calculation of two sets of key correlation weights: the "first correlation weight" quantifies the impact of each power harmonic data point on various CPU performance parameters (such as the contribution to the decrease in computing speed); the "second correlation weight" quantifies the impact of each power harmonic data point on various server temperature data (such as the contribution to the core temperature rise).
[0066] The calculated weight values form a correlation weight matrix, a mathematical table where rows may represent different harmonic components, columns represent different performance or temperature indicators, and the values at the intersections represent the quantified influence of that harmonic on that indicator. This matrix is key to transforming chaotic correlations into interpretable and comparable numbers.
[0067] Step S303: Based on the correlation weight matrix, construct the mapping relationship between the power supply harmonic content range and the server CPU performance attenuation range, as well as the mapping relationship between the power supply harmonic content range and the server temperature rise range, and generate the harmonic influence baseline curve.
[0068] The correlation weight matrix reveals the microscopic influence of parameters on parameters, but operations and maintenance personnel need a macroscopic guideline that can be directly compared: that is, when the total harmonic distortion reaches a certain level, how much might my server's CPU performance drop? How much might the temperature rise? This step is precisely to complete this conversion.
[0069] Specifically, based on a weight matrix and combined with a large amount of data distribution, the system divides continuous harmonic content (e.g., THD from 0% to 10%) into several representative content intervals (e.g., 0%-3% is excellent, 3%-5% is noteworthy, 5%-8% is a warning, and above 8% is severe). For each interval, the model integrates all sample data within that interval, as well as the weights of each harmonic component, to calculate the corresponding server CPU performance degradation range (e.g., the percentage range of decrease in computing speed) and server temperature rise range (e.g., the range of increase in core temperature by degrees Celsius). These two sets of "interval-amplitude / range" correspondences are visualized as the initial harmonic impact baseline curve. This curve (or a set of corresponding tables) is the culmination of the system's intelligence; it makes the output of complex mathematical models intuitive and operable, becoming the absolute benchmark for real-time cause determination in subsequent steps.
[0070] In the above implementation, a pre-trained model is invoked to achieve rapid startup. By calculating the correlation weight matrix, the complex nonlinear relationship between harmonics and server status is quantified in depth, thereby constructing an intuitive and usable harmonic impact benchmark curve, transforming the black-box output of machine learning into a white-box decision-making tool for operation and maintenance.
[0071] Reference Figure 4 As one implementation of step S104, the steps of real-time monitoring of server status data, and when the server status data exceeds a preset warning threshold, calling the harmonic influence reference curve to determine the cause and generating an anomaly determination report include: Step S401: Monitor server status data in real time. When the server CPU performance data or server temperature data exceeds the preset warning threshold, obtain the power harmonic data at the current moment. The system is equipped with preset warning thresholds based on server design specifications and business continuity requirements (such as a CPU core temperature of 85°C or a 10% decrease in processing speed compared to the baseline). Real-time monitoring means that the system continuously compares the latest collected CPU performance parameters (such as instruction latency and utilization) and server temperature data (such as core temperature and air intake temperature) with these thresholds at a high frequency (such as every minute). Once any parameter exceeds the threshold, it indicates that the server's operating status has entered an "abnormal" or "sub-healthy" alarm state.
[0072] At this point, the system immediately responds and acquires the power harmonic data at the current moment. Through a unified timestamp system, it ensures that the acquired harmonic data (such as total harmonic distortion rate and the content of each harmonic) and the server status data that triggered the alarm are observations at the same time. This provides the initial, time-aligned evidence necessary for subsequent causal analysis to determine whether the anomaly was caused by a power supply problem.
[0073] Step S402: Call the harmonic influence reference curve and extract the server status data influence range corresponding to the power harmonic data at the current moment; Among them, the harmonic impact baseline curve is a quantitative model that was previously trained on massive historical data through machine learning. It is essentially a knowledge graph or mapping function that describes "the possible range of server performance degradation and temperature rise under a given harmonic content".
[0074] In this embodiment, once the system obtains the specific harmonic data at the current moment (e.g., Total Harmonic Distortion (THD) = 6.5%), it invokes this baseline curve. The extraction operation involves finding or calculating the corresponding server status data impact range on the curve based on the current harmonic value. For example, the curve might indicate that when THD is in the range of 6.0%-7.0%, historical data shows a 90% probability of causing a decrease in CPU processing speed of 8% to 15% and an increase in core temperature of 4°C to 8°C. This "decrease of 8%-15%" and "increase of 4-8°C" constitute a quantified impact range.
[0075] The significance of this step is that it is not based on a single fixed threshold, but on a probability distribution-based interval prediction, providing a dynamic and precisely matched expected range of impact for subsequent judgments.
[0076] Step S403: Compare the current server status data change range with the range of the affected area, and analyze whether the server status data change trend is consistent with the predicted trend of the harmonic influence baseline curve. Step S404: Determine whether the change in the current server status data is within the influence range and whether the trend of the server status data change is consistent with the predicted trend of the harmonic influence baseline curve; if yes, proceed to step S405; if no, proceed to step S406. Specifically, the first step is to compare the magnitude of change: calculate the actual change in the current abnormal server parameters relative to its normal baseline value (e.g., the CPU processing speed actually decreased by 12%, and the core temperature actually increased by 6°C), and compare it with the impact range extracted in the previous step (predicted decrease of 8%-15%, predicted increase of 4-8°C). If the actual value falls within the predicted range (e.g., 12% is within 8%-15%, 6°C is within 4-8°C), then this constitutes the first layer of supporting evidence.
[0077] Secondly, a consistency analysis of server status data change trends is performed: the system analyzes the trajectory of server status data (such as temperature) changes within a time window before the anomaly occurs (e.g., whether it's a sudden jump or a gradual increase synchronized with harmonic content), and compares it with the typical harmonic impact pattern revealed by the baseline curve (e.g., harmonic-induced temperature rise is usually gradual and cumulative). Trend analysis effectively eliminates false positive alarms caused by instantaneous external interference (such as a spike in CPU usage due to an occasional high-load process) or momentary sensor malfunctions. Only when the actual change magnitude matches the predicted range, and the actual change trend also coincides with the typical predicted trend of harmonic effects, can a strong double confirmation be formed.
[0078] Step S405: Determine that the influence is dominated by harmonics; Step S406: Determine that other factors are influencing the outcome; If the comparison results simultaneously meet both of the above conditions, the system determines that the current server anomaly is mainly caused by harmonics. This means that data analysis shows that the current harmonic level is sufficient, and its mode of action is highly consistent with the observed server state deterioration. Therefore, there is reason to believe that harmonics are the main driving factor of this anomaly.
[0079] Otherwise, if any of the above conditions are not met (e.g., the amplitude exceeds the prediction range, or the trend is a sudden, slow accumulation pattern that does not conform to the harmonic effects), it is determined to be due to other factors. This clearly excludes the root cause of the problem from the power supply quality area, guiding maintenance personnel to investigate other possible causes, such as server local cooling system failure (e.g., fan stoppage), CPU overload caused by applications, or partial failure of the data center air conditioning, etc.
[0080] Step S407: Generate an anomaly determination report, which includes the determination result, current power harmonic data, and the magnitude of changes in server status data.
[0081] The anomaly determination report is a comprehensive document that includes the determination results (harmonic dominance / other factors), current power supply harmonic data details (such as specific THD values and harmonic spectra), and server status data change magnitudes (such as specific percentage of performance degradation, specific degree of temperature rise, and trend charts).
[0082] It's important to note that the value of this report lies in the following: For cases determined to be "harmonic-dominated," it provides the operations and maintenance team with a complete chain of evidence demonstrating the problem ("The current harmonic level is X, which typically leads to an impact in the Y range, and what is actually observed is Z, with matching trends"), greatly enhancing the persuasiveness of the warning and providing direct decision-making data support for subsequent remedial measures (such as installing filters). For cases determined to be "other factors," the detailed data in the report (such as normal harmonic data but a surge in temperature) can also help operations and maintenance personnel quickly troubleshoot power supply issues, focusing on server hardware or local environment faults.
[0083] In the above implementation, efficient triggering is achieved based on a preset warning threshold. Real-time data is input into a historical knowledge model by calling the harmonic impact benchmark curve. Rigorous hypothesis testing is performed through a dual comparison of amplitude and trend. Finally, a judgment result is output based on clear rules, and a comprehensive anomaly judgment report is generated. This technical solution not only significantly improves fault location efficiency and reduces blind troubleshooting by maintenance personnel, but more importantly, by effectively distinguishing between anomalies caused by harmonics and those not caused by harmonics, it can precisely guide limited maintenance resources towards the correct handling direction. This fundamentally prevents server performance degradation and downtime risks caused by power quality issues, ensuring the ultra-high reliability and stability of data center operations.
[0084] Reference Figure 5 As one implementation of step S105, the step of generating a graded early warning report and pushing it to a designated terminal based on the anomaly determination report includes: Step S501: Analyze the anomaly determination report and extract the determination results, current power harmonic data, and server status data change amplitude. The determination of whether the problem is "dominated by harmonics" or "affected by other factors" is crucial, as it determines the fundamental nature of the subsequent warning and the direction of the response. Secondly, the current power supply harmonic data, including the specific harmonic parameter values that triggered the determination, such as the precise value of the total harmonic distortion (THD) and the significantly exceeding harmonic current / voltage content, provides direct evidence for locating power quality issues. Finally, the server status data variation, specifically the actual monitored CPU performance degradation (e.g., percentage decrease) and temperature increase, quantifies the severity of the abnormal event.
[0085] Step S502: Based on the judgment result and the change range of server status data, match the preset graded early warning rules to determine the early warning level; The system has a pre-defined set of tiered early warning rules. These rules, which reflect the business logic, clearly define the action levels corresponding to different combinations of "judgment results" and different "degrees of server status data anomalies." For example, the rules might be defined as follows: if the judgment is "harmonic dominance," and the CPU performance decreases by 5%-10% or the temperature slightly exceeds the threshold, then a "Level 1 Warning" (informative attention) is triggered; if the performance decreases by 10%-20% or the temperature continues to exceed the limit, then a "Level 2 Warning" (planned intervention required) is triggered; and if the performance decreases by more than 20% or the temperature approaches the danger threshold, then a "Level 3 Warning" (emergency response required) is triggered.
[0086] In this embodiment, the matching process is an automated rule engine execution process. The system takes the extracted "judgment result" and the quantified "abnormality level" as input and compares them one by one with the conditions in the rule base. Determining the warning level is the output of this step. It is not a simple label, but triggers a whole set of subsequent processing logic for the corresponding level, including the report content template, push channel, notification scope, and expected response time.
[0087] Step S503: Based on the warning level, generate a graded warning report containing handling suggestions and send it to the designated terminal through the push interface.
[0088] The tiered early warning reports are generated in a customized manner based on the warning level. For a Level 1 warning, the report may focus on a detailed breakdown of power harmonic data and a brief presentation of the current server status data changes, along with recommendations for enhanced monitoring. For Level 2 and 3 warnings, the reports will include more detailed data, particularly server status data change curves (showing the trend of parameter deterioration over time), and core handling recommendations.
[0089] It should be noted that the handling suggestions are derived from the built-in operation and maintenance knowledge base. Based on the specific harmonic exceedance pattern (e.g., predominantly 5th and 7th harmonics) and the type of server anomaly, targeted remediation measures can be recommended. For example, suggestions such as "Install an active power filter for 5th and 7th harmonics in the P1 column cabinet" or "Immediately check if the air conditioning supply in cabinet A03 is functioning properly" are provided. The report generation process is an automated operation of template filling and data binding. The final tiered early warning report is a comprehensive work instruction integrating anomaly snapshots, trend analysis, root cause diagnosis, and action guidelines. Its information density and the strength of the action guidelines are positively correlated with the warning level.
[0090] Subsequently, the system automatically executes a piece of code that sends out the formatted hierarchical early warning report as a data payload through a pre-configured communication protocol (such as SMTP for email, SMS gateway for SMS, RESTful API for operation and maintenance platforms or collaborative tools such as WeChat / DingTalk).
[0091] In some embodiments, the setting of designated terminals is key to accurate push notifications, which are linked to the warning level: Level 1 warnings may only be pushed to the mobile applications of front-line maintenance personnel in the relevant area; Level 2 warnings will also be pushed to the email of the maintenance supervisor and the data center monitoring screen; Level 3 warnings may be additionally broadcast via SMS, automated voice calls, or even enterprise communication groups to ensure that the information is perceived by all relevant personnel as soon as possible.
[0092] The above implementation constitutes a highly automated, strategy-driven, and precisely targeted early warning information generation and distribution system. It achieves intelligent and standardized response strategies through tiered early warning rule matching, provides action guidelines with actionable knowledge matching the severity of the event through tiered report generation, and finally ensures that instructions are accurately and promptly delivered to the responsible parties through interface call push.
[0093] In practical applications, this technical solution transforms anomaly detection reports into dispatch orders that drive different levels and roles of maintenance personnel to take differentiated actions. This not only solves the problems of vague and indiscriminate monitoring and alarm information in traditional monitoring, but also ensures a seamless connection from intelligent system detection to efficient on-site handling through built-in handling suggestions. This shortens the average fault response and repair time and improves the accuracy and collaboration of data center operations and maintenance.
[0094] Reference Figure 6 As one implementation of step S106, the steps of recording operation and maintenance data based on the graded early warning report, optimizing the parameters of the correlation analysis model, and updating the harmonic impact baseline curve include: Step S601: Receive the graded early warning report, extract the early warning level, the magnitude of changes in server status data, and handling suggestions; Among them, the warning level defines the severity of the incident and the urgent priority of the handling background; the abnormal server status data, namely the specific deterioration values of CPU performance and temperature that triggered the warning recorded in the report, constitute the baseline status that needs to be verified by the "governance effect"; the handling suggestion is the action plan recommended by the system for this specific anomaly based on its knowledge base (e.g., "install an active filter on the P1 column cabinet").
[0095] Step S602: Based on the handling suggestions, obtain the handling operation data performed by the operation and maintenance personnel and the change range of server status data after the handling, generate an operation and maintenance handling record dataset and synchronize it to the historical database; The system does not automatically execute physical governance actions; instead, it guides and records manual operations. Based on handling suggestions extracted from early warning reports, it provides a structured interface (such as an operations and maintenance work order form) for operations and maintenance personnel to record the actual handling operations they perform.
[0096] Specifically, the maintenance and repair record dataset includes, but is not limited to: whether maintenance personnel followed the recommendations (e.g., installing the recommended filter model), the precise execution time, the specific location of the operation, and the equipment parameters used. More importantly, after the repair operation is completed and a period of stable operation has elapsed, the system will again collect the changes in server status data after the repair. This includes the power harmonic content after the repair (e.g., the total harmonic distortion rate decreased from 6.2% to 3.5%), as well as the corresponding recovery of server CPU performance and temperature (e.g., the core temperature dropped from 86℃ to 78℃, and the computing speed returned to normal).
[0097] Furthermore, the generated operational handling record dataset can be added to or updated as a new record with complete timestamps and event tags via standard data interfaces (such as database API calls). This historical database typically employs a structure suitable for storing time-series and event data; it not only stores the raw monitoring data but, more importantly, these "handling-feedback" records with clear causal markers. This operation allows the experience gained from a single operational event to be solidified and preserved, and then aggregated with data from all other similar events throughout history.
[0098] Step S603: Invoke the correlation analysis model, load the operation and maintenance handling record dataset from the historical database, and optimize the correlation analysis model parameters; The system periodically, or when new data accumulates to a certain scale, invokes the correlation analysis model previously used to generate the harmonic impact baseline curve. Unlike the initial training, the core data loaded for this optimization is a dataset of operation and maintenance records from the historical database. This data not only contains the correlation between harmonics and server status, but also clearly defines how a specific "intervention" (such as installing a certain type of filter) alters this correlation.
[0099] In this embodiment, the model optimization process essentially uses these new data with feedback results as training samples to run machine learning algorithms (such as gradient descent) and adjust the parameters within the model (e.g., adjusting the splitting conditions of decision tree nodes in a random forest, or the connection weights of neurons in a neural network). The goal of optimization is to make the model's prediction of the "post-treatment state" more closely resemble the actual "post-treatment parameter changes" recorded in the database. This means that after learning new feedback, the model updates its quantitative understanding of "what harmonic mode leads to what impact" and "what governance measures might produce what effect."
[0100] The logic of this step lies in encapsulating and outputting the adjustment results, such as internal weights and coefficients calculated during model optimization, so that they can replace the old parameters and be applied to the subsequent real-time analysis and early warning process. The model optimization calculation process generates a series of new mathematical coefficients, weight matrices, and threshold sets, which are the updated correlation analysis model parameters. This generation action means that the system exports these new parameters from the training process, formats and versions them, and these new parameters are then prepared to be used to update the online correlation analysis model instance.
[0101] Step S604: Update the correlation weight matrix and harmonic influence baseline curve based on the optimized correlation analysis model.
[0102] Once the optimized correlation analysis model is deployed, the correlation weight matrix calculated based on it and the final harmonic impact baseline curve will also be updated. When the system subsequently performs correlation analysis, cause determination, and other steps, the harmonic impact baseline curve used will be a new version defined by this new set of parameters. The new model incorporates experience learned from one or more recent successful (or failed) maintenance and repair operations, resulting in a more accurate grasp of the patterns of harmonic impact on servers, more reliable identification of the root causes of anomalies, and potentially more targeted recommended solutions.
[0103] Understandably, as the model parameters are updated, the entire system is able to adapt to changes in the data center, such as responding more accurately to newly added server types and new load patterns. This represents a leap from "training based on historical data" to "continuous learning during operation," ensuring the long-term effectiveness and accuracy of the early warning system.
[0104] In the above implementation, the handling operations and results are recorded to digitize human experience. This data is then synchronized to a historical database for continuous knowledge accumulation. Finally, by calling the model and loading new data to optimize parameters, the system can learn from each operational practice, verify hypotheses, and correct its understanding. This closed loop ensures that the system's core analysis model does not stagnate but evolves continuously with data center equipment updates, load changes, and the implementation of governance measures, resulting in increasingly accurate harmonic impact baseline curves. This technical solution endows the solution with long-term vitality and scenario adaptability, fundamentally guaranteeing that the accuracy of early warnings continuously improves over time. It provides the technical support for enabling data center operations to move from "experience-driven" to "data intelligence-driven" and possess "continuous learning" capabilities.
[0105] Reference Figure 7 As a further implementation of the intelligent server status early warning method, the early warning method also includes: Step S701: Collect server load data in real time and generate a server load dataset with timestamps; In the original solution, when the server was in an abnormal state, the system could only attribute it to "other factors" after determining that it was caused by non-harmonic interference. The maintenance personnel still had to manually troubleshoot among various possibilities such as heat dissipation failure and application overload.
[0106] Therefore, this step involves collecting server load data in real time through server management interfaces (such as IPMI, BMC, or operating system agents). This includes a series of indicators reflecting the supply and demand of server computing resources, such as CPU utilization, memory usage, disk I / O rate, network traffic, and task queue length. By binding timestamps to these data, they are placed on the same time coordinate system as harmonic, performance, and temperature data.
[0107] Understandably, the underlying logic for generating server load datasets lies in the fact that server load is the direct and primary internal factor driving CPU operations and heat generation. High load inevitably leads to increased CPU utilization and power consumption, which can then cause performance queuing delays and temperature increases. By introducing this dataset, the system gains a crucial key to distinguishing between "temperature rise / performance changes caused by normal working load" and "additional temperature rise / performance degradation caused by abnormal power supply quality (harmonics)." For example, if a server experiences extremely high CPU utilization and task queues at high temperatures, the anomaly is likely driven by the load itself; conversely, if the load is stable but the temperature spikes, the suspicion of harmonics or heat dissipation issues increases significantly.
[0108] Step S702: Perform timestamp alignment between the server load dataset and the standardized analysis dataset to generate a time-synchronized dataset; Although each dataset has its own timestamp, there may be non-systematic offsets in the millisecond to second range between load data from the server operating system and harmonic data from power monitoring instruments due to sampling frequency, slight differences in system clock, and network latency.
[0109] In this embodiment, the timestamp alignment operation forcibly aligns these heterogeneous data streams onto a unified timeline using high-precision algorithms (such as matching based on the nearest neighbor timestamp or interpolating and resampling higher-frequency data to lower-frequency data time points). A time-synchronized dataset means that in this dataset, the load value, harmonic value, CPU performance value, and temperature value contained in any record row strictly correspond to the same precisely defined moment.
[0110] Step S703: Perform data integration processing on the time synchronization dataset to generate an extended analysis dataset; The data integration and processing takes over the time-synchronized data and performs a series of standardized operations: First, format unification and unit calibration are performed to ensure that load data (such as CPU utilization percentage), harmonic data (such as THD percentage), performance data (such as instruction latency in nanoseconds), and temperature data (degrees Celsius) are in a standardized format that allows for mutual computation.
[0111] Secondly, missing value handling is required. For individual data points that are missing due to synchronization or acquisition failure, reasonable imputation methods (such as the average of the previous and next time points or correlation-based prediction imputation) are used to ensure the continuity of the data sequence and avoid information interruption during model training.
[0112] More crucially, feature engineering extraction involves creating derived features that are more predictive, such as calculating the rolling average of load parameters to smooth out instantaneous spikes, calculating the rate of change (slope) of load within a short window to characterize its growth trend, or constructing the interaction term (product) between load and harmonics to preliminarily detect their synergistic effect.
[0113] After these processes, the final extended analysis dataset is no longer just a simple parallel arrangement of four independent data streams, but a comprehensive data entity that is internally consistent and can more fully depict how "external power quality" and "internal workload" work together to affect "server operating status".
[0114] Step S704: Call the correlation analysis model to perform machine learning-driven correlation analysis on the extended analysis dataset, calculate the correlation weights of server load parameters, power harmonic data and server status data, and generate a comprehensive influence baseline curve. The original scheme mainly focuses on the unilateral impact of harmonics, while the model in this step (such as gradient boosting tree, deep neural network) uses the aforementioned extended analysis dataset as training samples, in which server load parameters and harmonic parameters are used as input features, and server status data is used as the prediction target.
[0115] Specifically, the machine learning-driven correlation analysis process learns autonomously and outputs two sets of key information: first, correlation weights, which precisely quantify the independent contribution of each load feature (such as CPU utilization) and each harmonic feature (such as the 5th harmonic current) to each state index (such as core temperature); second, capturing the interaction weights between features, for example revealing whether the common effect of "high load" and "high harmonics" on temperature rise is a simple addition or a "1+1>2" amplification effect when they coexist.
[0116] Ultimately, all these learned complex relationships are summarized and expressed as one or more comprehensive influence baseline curves (or high-dimensional response surfaces). This curve is a powerful predictive tool that can answer complex questions such as, "What is the expected range of CPU core temperature under conditions of 70% CPU utilization and 4% total harmonic distortion?" or "To ensure that the temperature does not exceed the safe threshold, under a given current load, what harmonic content must be kept below?" It makes explicit and tool-based the implicit, multi-factor coupled influence patterns.
[0117] Step S705: Dynamically adjust the preset early warning threshold based on the comprehensive impact baseline curve; The logic of this step is to enable the system's alarm trigger line to adaptively float according to the changes in the server's real-time workload, thereby significantly reducing invalid alarms (false alarms) under high-load normal operation scenarios and enhancing the detection sensitivity of subtle anomalies (such as pure harmonic effects) under low load (reducing missed alarms).
[0118] Understandably, the original preset warning thresholds (such as a fixed CPU temperature of 85°C) are static and cannot distinguish the essential difference between a server reaching 85°C in an idle state and reaching 85°C under full load computing state (the former is extremely abnormal, while the latter may be within expectations).
[0119] This step utilizes the knowledge contained in the comprehensive influence baseline curve to dynamically calculate the reasonable expected range of server status data under the current load level based on real-time collected server load data. The system then dynamically adjusts the warning threshold accordingly. For example, if the model predicts that the reasonable upper limit of CPU core temperature is 88℃ under 80% CPU load, the system will raise the temporary threshold from 85℃ to 88℃; conversely, when the load is only 10%, the model predicts that the normal temperature should be very low, and the system may lower the threshold to 82℃ to more sensitively capture abnormal temperature rises that may be caused by harmonics or other issues. This mechanism ensures that the warning signal always targets "anomalies exceeding the reasonable expected range of the current workload," rather than "exceeding a certain absolute fixed value," which greatly improves the specificity of the alarms and the trust of operations and maintenance personnel in the alarms.
[0120] Step S706: Monitor server status data in real time and extract the trend characteristics of current server load data. While monitoring status parameters in real time, the system performs sliding window analysis on continuous server load data (such as CPU utilization sequences) and extracts trend characteristics. This includes using linear regression to calculate the slope, calculating the derivative of a moving average, or performing simple spectral analysis to quantify the dynamic behavior of the load. These trend characteristics may include: short-term rising / falling slopes (indicating whether the load is increasing or decreasing rapidly), fluctuation intensity (indicating the stability of the load), and peak persistence patterns (indicating whether high load is a momentary spike or a sustained plateau).
[0121] It's important to note that the logical necessity of this step lies in the fact that different root causes often correspond to different load change patterns. For example, a planned batch computing task might cause the load to rise slowly and remain at a high level (gradual plateau), with the resulting temperature rise being expected; while an unexpected software infinite loop or resource contention might cause the load to spike to 100% without warning (instantaneous spike), potentially triggering immediate performance issues; performance degradation and temperature rise caused by harmonic issues might manifest as a slow, continuous deterioration trend under relatively stable load conditions. By extracting these trend characteristics, the goal is to transform the load from a static scalar into a vector containing its dynamic behavior, thereby laying a more solid foundation for distinguishing different types of anomalies.
[0122] Step S707: When the server status data exceeds the adjusted preset warning threshold, the cause is determined by combining the comprehensive impact baseline curve and the trend characteristics, and an optimized anomaly determination report is generated.
[0123] The logic of this step lies in comprehensively utilizing three types of information—dynamic threshold, quantitative influence curve, and load trend characteristics—to diagnose abnormalities exceeding the standard. This allows for a clear distinction between "load-dominated abnormalities," "harmonic-dominated abnormalities," "abnormalities caused by a combination of load and harmonic effects," and "other (such as heat dissipation) fault-related abnormalities."
[0124] Specifically, when the state parameters trigger the adjusted dynamic threshold (which itself filters out normal changes caused by pure high load), the system initiates enhanced diagnostics. First, it calls the harmonic impact baseline curve to check whether the current harmonic level is sufficient to explain the observed state deterioration on its own (amplitude and trend comparison). At the same time, it analyzes the characteristics of the changing trends in depth: if the load shows a short-term sharp upward trend and is highly synchronized with the time of state deterioration, it strongly points to "load dominance"; if the load is stable but the harmonics exceed the standard and match the baseline curve prediction, it points to "harmonic dominance"; if the load increases slowly and the harmonics also increase slowly, and the state deterioration is the superposition of the two, it may be judged as "mixed influence"; if neither of them shows obvious abnormalities, it points to "other faults" such as the cooling system.
[0125] Ultimately, the generated optimized anomaly determination report not only includes the determination result, but also lists in detail the chain of evidence supporting the conclusion: including the basis for dynamic threshold adjustment, analysis charts of load trends, comparison of harmonic data with the baseline curve, and probability assessments of various possibilities.
[0126] In the above implementation, the system can clearly distinguish whether the server's abnormal status stems from inherent changes in its internal workload, interference from external power supply quality (harmonics), or a combination of both, thus resolving the ambiguity in the "other factors" determination in the original solution. By dynamically adjusting the warning threshold, the system intelligently adapts to the natural fluctuations in server load, reducing the false alarm rate and enhancing the sensitivity to detect hidden problems. Combined with the analysis of load change trend characteristics, the system can not only diagnose existing anomalies but also perceive emerging risk patterns. The resulting optimized anomaly determination report provides unprecedented root cause transparency and action guidance.
[0127] In practical applications, this technical solution elevates data center operations and maintenance from responding to isolated symptoms to a deeper understanding and precise control of the multivariate interactions in complex systems. It achieves a substantial improvement in the scientific and intelligent level of operations and maintenance decision-making, providing core technical support for ensuring the stability and optimal energy efficiency of high-density, high-complexity data center infrastructure.
[0128] This application also discloses an intelligent early warning system for server status based on a harmonic correlation model.
[0129] A server status intelligent early warning system based on a harmonic correlation model, specifically comprising: The multi-source data synchronous acquisition module is used to synchronously acquire power harmonic data and server status data from the data center, generating a raw multi-source dataset with timestamps; among which, the server status data includes server CPU performance data and server temperature data. The data processing module is used to perform timestamp alignment and preprocessing on the original multi-source datasets to generate standardized analysis datasets; The baseline curve generation module is used to call the correlation analysis model, perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power supply harmonic data and server status data, and generate a harmonic influence baseline curve. The anomaly detection module is used to monitor server status data in real time. When the server status data exceeds the preset warning threshold, it calls the harmonic influence benchmark curve to determine the cause and generates an anomaly detection report. The tiered early warning module is used to generate tiered early warning reports based on anomaly determination reports and push them to designated terminals; The feedback optimization module is used to record operation and maintenance data based on hierarchical early warning reports and optimize the parameters of the correlation analysis model.
[0130] The server status intelligent early warning system based on the harmonic correlation model in this application embodiment can implement any of the above methods, and the specific working process of each module in the system can refer to the corresponding process in the above method embodiments.
[0131] In the several embodiments provided in this application, it should be understood that the provided methods and systems can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for example, the division of a certain module is merely a logical functional division, and in actual implementation there may be other division methods, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0132] This application also discloses a computer device.
[0133] A computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement a server status intelligent early warning method based on a harmonic correlation model as described above.
[0134] This application also discloses a computer-readable storage medium.
[0135] A computer-readable storage medium storing a computer program that can be loaded by a processor and executed as described above in any of the server status intelligent early warning methods based on a harmonic correlation model.
[0136] The computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device; the program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0137] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A server status intelligent early warning method based on a harmonic correlation model, characterized in that, The method includes: Synchronously collect power harmonic data and server status data from the data center to generate a raw multi-source dataset with timestamps; the server status data includes server CPU performance data and server temperature data. The original multi-source dataset is timestamped and preprocessed to generate a standardized analysis dataset. The correlation analysis model is invoked to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power harmonic data and server status data, and generate a harmonic impact baseline curve. The server status data is monitored in real time. When the server status data exceeds the preset warning threshold, the harmonic influence reference curve is called to determine the cause and generate an anomaly determination report. Based on the anomaly determination report, a tiered early warning report is generated and pushed to the designated terminal; Based on the tiered early warning report, operation and maintenance data are recorded, and the parameters of the correlation analysis model are optimized to update the harmonic impact baseline curve.
2. The intelligent early warning method for server status based on a harmonic correlation model according to claim 1, characterized in that, The steps for generating a standardized analysis dataset by timestamping and preprocessing the original multi-source dataset include: Based on the timestamps of the original multi-source dataset, power harmonic data, server CPU performance data, and server temperature data are aligned and integrated to generate an aligned and integrated dataset. The aligned and integrated dataset is cleaned to remove outliers and generate a cleaned dataset. The cleaned dataset is formatted and the data units are standardized to a preset standard to generate a standardized dataset. The standardized dataset is normalized to eliminate the differences in the numerical magnitude of different parameters, thereby generating a standardized analysis dataset.
3. The intelligent early warning method for server status based on a harmonic correlation model according to claim 1, characterized in that, The steps of calling the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculating the correlation weights between power supply harmonic data and server status data, and generating a harmonic impact baseline curve include: Call the pre-trained correlation analysis model and input the power harmonic data and server status data from the standardized analysis dataset; The first correlation weight between power supply harmonic data and server CPU performance data, and the second correlation weight between power supply harmonic data and server temperature data are calculated using machine learning algorithms to generate a correlation weight matrix. Based on the aforementioned correlation weight matrix, a mapping relationship is constructed between the power supply harmonic content range and the server CPU performance attenuation range, as well as a mapping relationship between the power supply harmonic content range and the server temperature rise range, generating a harmonic influence baseline curve.
4. The intelligent early warning method for server status based on a harmonic correlation model according to claim 1, characterized in that, The steps of real-time monitoring of the server status data, and generating an anomaly determination report by calling the harmonic influence reference curve to determine the cause when the server status data exceeds a preset warning threshold, include: Real-time monitoring of server status data; when server CPU performance data or server temperature data exceeds the preset warning threshold, obtain the power harmonic data at the current moment. Call the harmonic impact baseline curve to extract the server status data impact range corresponding to the current moment's power harmonic data; Compare the current server status data change range with the range of the influence interval, and analyze whether the server status data change trend is consistent with the predicted trend of the harmonic influence benchmark curve; If the change in the current server status data is within the influence range and the trend of the server status data change is consistent with the predicted trend of the harmonic influence baseline curve, then it is determined to be a harmonic-dominated influence; otherwise, it is determined to be influenced by other factors. Generate an anomaly assessment report, including the assessment results, current power harmonic data, and the magnitude of changes in server status data.
5. The intelligent early warning method for server status based on a harmonic correlation model according to claim 4, characterized in that, The steps of generating a tiered early warning report and pushing it to a designated terminal based on the anomaly determination report include: Analyze the anomaly determination report to extract the determination result, current power harmonic data, and server status data change amplitude; Based on the judgment result and the magnitude of changes in server status data, a preset graded early warning rule is matched to determine the early warning level; Based on the warning level, a graded warning report containing handling suggestions is generated and sent to the designated terminal through a push interface.
6. The intelligent early warning method for server status based on a harmonic correlation model according to claim 3, characterized in that, Based on the tiered early warning report, the steps of recording operation and maintenance data, optimizing the parameters of the correlation analysis model, and updating the harmonic impact baseline curve include: Receive tiered early warning reports, extract the early warning level, the magnitude of changes in server status data, and handling suggestions; Based on the proposed handling suggestions, obtain the handling operation data performed by the operation and maintenance personnel and the change range of server status data after the handling, generate an operation and maintenance handling record dataset and synchronize it to the historical database; Call the correlation analysis model, load the operation and maintenance record dataset from the historical database, and optimize the correlation analysis model parameters; Based on the optimized correlation analysis model, the correlation weight matrix and harmonic influence baseline curve are updated.
7. The intelligent early warning method for server status based on a harmonic correlation model according to claim 4, characterized in that, The early warning method also includes: Real-time collection of server load data, generating a server load dataset with timestamps; The server load dataset is timestamped and aligned with the standardized analysis dataset to generate a time-synchronized dataset. Perform data integration processing on the time-synchronized dataset to generate an extended analysis dataset; The association analysis model is invoked to perform machine learning-driven association analysis on the extended analysis dataset, calculate the association weights of server load parameters, power harmonic data and server status data, and generate a comprehensive influence baseline curve. Based on the comprehensive impact benchmark curve, the preset early warning threshold is dynamically adjusted; The server status data is monitored in real time, and the changing trend characteristics of the current server load data are extracted. When the server status data exceeds the adjusted preset warning threshold, the cause is determined by combining the comprehensive impact benchmark curve and the trend characteristics, and an optimized anomaly determination report is generated.
8. A server status intelligent early warning system based on a harmonic correlation model, characterized in that, The system includes: The multi-source data synchronous acquisition module is used to synchronously acquire power harmonic data and server status data from the data center, generating a raw multi-source dataset with timestamps; among which, the server status data includes server CPU performance data and server temperature data. The data processing module is used to perform timestamp alignment and preprocessing on the original multi-source datasets to generate standardized analysis datasets; The baseline curve generation module is used to call the correlation analysis model to perform machine learning-driven correlation analysis on the standardized analysis dataset, calculate the correlation weight between power harmonic data and server status data, and generate a harmonic influence baseline curve. The anomaly detection module is used to monitor the server status data in real time. When the server status data exceeds the preset warning threshold, the module calls the harmonic influence reference curve to determine the cause and generates an anomaly detection report. The graded early warning module is used to generate a graded early warning report based on the anomaly determination report and push it to the designated terminal; The feedback optimization module is used to record operation and maintenance data based on the hierarchical early warning report and optimize the parameters of the correlation analysis model.
9. A computer device, characterized in that: The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed as described in any one of claims 1 to 7.