A large computing power data center fault early warning intelligent response method
By deploying distributed sensor networks and deep learning models in high-performance data centers, and combining blockchain technology for multi-dimensional data collection and response strategy optimization, the shortcomings of existing fault warning systems have been addressed. This has enabled highly accurate fault warnings and rapid responses, ensuring the stability and business continuity of the data center.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO SIHONG ELECTRICAL APPLIANCE IND
- Filing Date
- 2026-04-10
- Publication Date
- 2026-05-29
AI Technical Summary
Existing fault early warning systems in high-performance computing data centers have shortcomings in data collection, early warning analysis, and response strategies. They cannot meet complex needs, resulting in high false alarm rates, high false negative rates, inaccurate response measures, difficulty in optimizing resource allocation, lack of cross-device linkage capabilities, susceptibility to secondary faults, and inability to adapt to fault sharing in multi-region deployments.
Multi-dimensional data collection is achieved using a distributed sensor network and server monitoring agent. Data preprocessing is performed using sliding window filtering and isolated forest algorithms. A fault prediction model based on deep learning is constructed, and a mapping database between fault types and response measures is established. Cross-device linkage is achieved through industrial Ethernet, and a cross-regional resource sharing platform is established. Intelligent load migration and environmental emergency control are adopted to monitor and optimize response strategies in real time.
It has achieved more precise fault early warning, intelligent response strategies, and coordinated equipment linkage, which has improved the operational stability and risk resistance of the data center, reduced operation and maintenance costs, and ensured the continuous operation of core businesses.
Smart Images

Figure CN122111798A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of alarm system technology, and in particular to an intelligent response method for early warning of faults in high-computing-power data centers. Background Technology
[0002] High-performance data centers, as the core infrastructure of the digital economy, bear critical tasks such as artificial intelligence training, large-scale data processing, and core business deployment. Their operational stability directly affects enterprise productivity, the continuity of social services, and data security. With the exponential growth in computing power demand, the scale of data center server clusters continues to expand, and hardware density, power consumption, and heat generation are rising simultaneously, significantly increasing the probability and scope of failures. Whether it's aging hardware components, uncontrolled data center environment, power fluctuations, or network congestion, all can trigger serious consequences such as server downtime and business interruptions, causing huge economic losses and reputational damage.
[0003] Existing fault early warning and response technologies have many shortcomings and are unable to meet the complex needs of high-performance data centers. At the data acquisition level, traditional systems often focus on single devices or single types of parameters, lacking comprehensive coverage of hardware operation, environmental status, power supply, and network load. Data acquisition frequencies are fixed and cannot be dynamically adjusted according to device importance and operating status, resulting in incomplete fault feature capture. In the early warning analysis stage, reliance on traditional threshold judgments and simple machine learning models is insufficient for extracting features from complex faults, is susceptible to data noise, and suffers from high false alarm and false negative rates. Furthermore, early warning thresholds are mostly fixed values, unable to adapt to dynamic scenarios such as seasonal changes and peak business periods.
[0004] In terms of response strategies and execution, existing methods offer relatively simplistic response measures, lacking precise matching with fault type, impact scope, and business priority. They often employ a one-size-fits-all approach, making it difficult to optimize resource allocation. Cross-device coordination is weak; servers, cooling systems, UPS power supplies, and other equipment operate independently, resulting in high latency in response command execution and a tendency for secondary failures due to uncoordinated actions. Furthermore, the lack of effective response effectiveness evaluation and model self-optimization mechanisms hinders the accumulation of fault handling experience, making it difficult to improve the accuracy of early warning responses over long-term operation. In addition, for data centers deployed in multiple regions, the lack of cross-regional resource scheduling and fault-sharing mechanisms means that a single regional failure can easily paralyze the entire business. These problems severely restrict the safe and stable operation of high-performance computing data centers. Summary of the Invention
[0005] The present invention proposes an intelligent response method for fault early warning in high-computing-power data centers to solve the problems mentioned in the prior art.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a smart response method for fault early warning in high-computing-power data centers, comprising the following steps: The data acquisition process involves deploying a distributed sensor network and server monitoring agents to collect hardware device operating parameters, data center environment parameters, power supply parameters, and network load data in real time. The data acquisition frequency is dynamically adjusted based on the importance of the equipment. In the preprocessing step, a sliding window filtering algorithm is used to reduce noise in the collected data. Normalization technology is used to unify the numerical range of parameters with different dimensions. The isolated forest algorithm is combined to identify outliers in the data, mark suspicious fault features, and generate a standardized dataset. The early warning analysis steps involve constructing a fault prediction model based on deep learning, inputting a standardized dataset for feature extraction and fault type matching, classifying early warning levels according to the probability of fault occurrence, scope of impact, and difficulty of recovery, and setting early warning thresholds. The strategy matching step involves establishing a mapping database between fault types and response plans, and matching corresponding response measures for different fault types. Cross-device linkage execution steps: Through the industrial Ethernet and device control interface, response commands are sent to the devices to realize the coordinated action of multiple devices, and the execution status is fed back in real time during the linkage process; The response effectiveness verification step involves continuously collecting relevant parameters after the response is executed, comparing the parameter changes before and after the fault is resolved, evaluating the effectiveness of the response measures, and triggering a secondary response if the fault is not eliminated, adjusting the response strategy or escalating the processing level. The logging and parameter update process records the fault occurrence time, fault type, warning level, response measures, execution process and processing results in detail, forming a fault handling log. Based on the log data, the parameters and response strategy library of the fault prediction model are optimized.
[0007] Furthermore, it also includes a dynamic quantification step for fault risk, using formulas. Calculate the fault risk value, where R is the overall fault risk value. The hardware loss weight is P, where P is the actual hardware runtime. For the design lifespan of hardware, The influence of temperature is weighted, where T is the current ambient temperature. For safe temperature threshold, This is the limiting temperature threshold. The load percentage is the weight, and L is the current device load rate. For maximum load rate, Here, W represents the power fluctuation weight, and W represents the actual voltage fluctuation amplitude. This is the standard voltage fluctuation threshold.
[0008] Furthermore, it also includes a dynamic early warning threshold calibration step. Based on historical fault data and real-time operating status, the early warning thresholds of each parameter are calibrated regularly. During the calibration process, the distribution pattern of parameters and the correlation with faults are analyzed. For parameters with frequent fluctuations, an adaptive threshold range is adopted, while for stable parameters, a fixed threshold is adopted. At the same time, the threshold range is adjusted in combination with seasonal changes and peak business scenarios. The calibration results are synchronized to the early warning analysis model in real time.
[0009] Furthermore, it also includes a rapid hardware fault location step, which uses blockchain-based device identity identification technology to assign a unique identifier to each hardware component, linking component production information, maintenance records and real-time operation logs to form a full lifecycle data chain; When a fault occurs, the feature vector output by the fault prediction model is combined with the hardware fault feature library to quickly locate the suspected faulty component. For clustered servers, a distributed fault diagnosis algorithm is used to narrow down the fault range to a specific slot or module by comparing the running data of adjacent nodes and analyzing the time series features.
[0010] Furthermore, the process includes intelligent load migration optimization steps. After a failure occurs, the resource scheduling center scans the remaining computing power, memory capacity, network bandwidth, and storage I / O performance of healthy nodes within the cluster in real time to select target servers that meet the operational requirements of the failed business. Based on virtual machine migration technology and container orchestration protocols, an incremental migration mode is adopted, prioritizing the migration of core business data and session states.
[0011] Furthermore, it also includes emergency control measures for the data center environment. In response to abnormal temperature faults, a distributed temperature sensor network is used to accurately locate hotspot areas, and the precision air conditioning system is linked to adopt a zone control strategy to adjust the air supply temperature and wind speed corresponding to the hotspot areas. At the same time, local cooling devices are activated to directly cool down the high-load server cluster.
[0012] Furthermore, it also includes a dynamic priority ranking step, using a formula. Determine the response priority, where P is the response priority coefficient, S is the number of devices affected by the fault, E is the importance coefficient of the fault's impact on business, D is the fault recovery difficulty coefficient, and C is the response execution cost coefficient.
[0013] Furthermore, it also includes power failure emergency response steps, real-time monitoring of mains voltage, frequency and phase changes, and immediate triggering of UPS power switching when a mains power interruption or parameters exceeding the safe range are detected, with switching time controlled in milliseconds; at the same time, the diesel generator emergency power supply process is started, and the generator fuel quantity, operating status and output parameters are monitored in real time. Based on load priority, the power supply strategy prioritizes the power supply to critical equipment such as database servers, switches, and monitoring systems. It monitors the remaining power and discharge rate of UPS batteries in real time, predicts the battery life based on battery health profiles, and automatically triggers a load shedding sequence when the battery life is insufficient, gradually cutting off unnecessary loads according to priority.
[0014] Furthermore, it also includes cross-regional linkage response steps. For high-performance data centers deployed in multiple regions, high-speed communication links and resource sharing platforms are established between regions, and the servers, storage, network and power resources of each region are virtualized into a global resource pool. When a major failure occurs in a single region and cannot be resolved independently, a resource request is automatically sent to the resource scheduling center of the adjacent region, specifying the required computing power, storage, and network bandwidth, etc. Based on the resource usage and fault impact range of each region, the resource sharing platform uses a latency-aware load distribution algorithm to migrate the business load of the faulty region to the nearest healthy region node with the lowest latency.
[0015] Furthermore, it also includes a model self-optimization iteration step, regularly iterating and optimizing the fault prediction model and response strategy library, and using reinforcement learning algorithms to train the model based on successful cases and failure experiences in the fault handling log; The model's feature weights and decision logic are continuously adjusted, with reward indicators being the accuracy of fault warnings, response time, and business recovery rate, and penalty indicators being the failure to issue timely warnings or respond to failures. Establish a fault simulation training library, generate diverse fault scenarios based on historical fault data, including single faults, compound faults, and cascading faults, and perform offline training and online fine-tuning of the model.
[0016] Compared with existing technologies, the beneficial effects of this invention are: This invention presents a high-performance data center fault early warning and intelligent response method that significantly improves the accuracy and foresight of early warnings. Multi-dimensional data collection covers parameters across all scenarios, including hardware, environment, power, and network, with dynamic adjustment of the collection frequency to ensure comprehensive capture of fault characteristics. A deep learning-based fault prediction model, combined with the isolated forest algorithm and multi-model fusion technology, effectively eliminates data noise and accurately identifies the characteristics of single, compound, and cascading faults, significantly reducing false alarm and false negative rates. Dynamic early warning threshold calibration and fault risk quantification mechanisms adapt to different operating scenarios, predicting fault risks in advance and achieving a shift from passive response to proactive early warning.
[0017] Response efficiency and targeting have been significantly optimized. The intelligent response strategy library establishes a precise mapping between fault types and response measures, dynamically prioritizing responses to ensure core business operations and optimize resource allocation. Cross-device collaboration, through industrial Ethernet and a unified control interface, enables coordinated actions of servers, cooling systems, UPS power supplies, and other devices, reducing response command execution latency and preventing secondary failures caused by uncoordinated actions. Specialized response measures such as intelligent load migration, environmental emergency control, and power failure emergency response target different fault types precisely, improving fault resolution efficiency.
[0018] Business continuity and operational convenience are comprehensively enhanced. A cross-regional collaborative response mechanism virtualizes resources from multiple regions into a global resource pool, enabling rapid migration and load sharing of business loads in faulty areas, preventing overall business paralysis caused by a single-region failure. A model self-optimization iteration mechanism continuously optimizes model parameters and response strategies based on fault handling logs, accumulating fault handling experience and improving early warning response performance after long-term operation. Rapid hardware fault location technology shortens fault investigation time, and response effect verification and secondary response mechanisms ensure thorough fault resolution, significantly reducing business interruption duration and data loss risk.
[0019] This invention, through innovation across the entire process, achieves more precise fault early warning, more intelligent response strategies, more collaborative equipment linkage, and more efficient operation and maintenance management. It effectively improves the operational stability and risk resistance of high-performance data centers, reduces operation and maintenance costs, provides reliable assurance for the continuous operation of core businesses, and is suitable for the application needs of large-scale, high-density, and multi-regional deployments of high-performance data centers. It has significant practical and economic value. Attached Figure Description
[0020] Figure 1 This is a schematic block diagram of an intelligent response method for fault early warning in a high-computing-power data center proposed in this invention. Figure 2 A bar chart comparing the accuracy of early warnings for different fault types; Figure 3 Line chart comparing the success rate of cross-regional load migration for different migration distances; Figure 4 A bar chart comparing response latency under different business loads. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0023] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.
[0024] Reference Figures 1 to 4 A method for intelligent response to fault early warning in high-computing-power data centers, comprising the following steps: The data acquisition process involves deploying a distributed sensor network and server monitoring agent to collect real-time hardware device operating parameters, data center environmental parameters, power supply parameters, and network load data. Hardware parameters include CPU load, memory usage, and hard disk I / O speed, while environmental parameters include temperature, humidity, and airflow speed. Power parameters include voltage, current, and power loss. The data acquisition frequency is dynamically adjusted based on the importance of the equipment. The preprocessing step employs a sliding window filtering algorithm to reduce noise in the collected data, normalization technology to unify the numerical range of parameters with different dimensions, and an isolated forest algorithm to identify outliers in the data, marking suspicious fault features and generating a standardized dataset. This supports collaboration between local preprocessing at edge nodes and centralized processing in the cloud. The early warning analysis step constructs a fault prediction model based on deep learning, inputs the standardized dataset for feature extraction and fault type matching, and classifies early warning levels according to the probability of fault occurrence, the scope of impact, and the difficulty of recovery. It also sets early warning thresholds for levels one, two, and three, with the early warning thresholds supporting dynamic optimization based on historical data. The strategy matching process establishes a mapping database between fault types and response solutions. For different fault types such as hardware failure, abnormal temperature, power fluctuation, and network congestion, corresponding response measures are matched, including load migration, equipment frequency reduction, emergency cooling, and power switching. The response strategy supports custom configuration and priority sorting. The cross-device linkage execution step sends response commands to devices such as server clusters, cooling systems, UPS power supplies, and network switches via industrial Ethernet and device control interfaces to achieve coordinated action of multiple devices. The execution status is fed back in real time during the linkage process to ensure that the response commands are accurately implemented. The response effect verification step continuously collects relevant parameters after the response is executed, compares the parameter changes before and after the fault is resolved, evaluates the effectiveness of the response measures, and triggers a secondary response if the fault is not eliminated, adjusting the response strategy or upgrading the processing level. The logging and parameter update process records the fault occurrence time, fault type, warning level, response measures, execution process and processing results in detail, forming a fault handling log. Based on the log data, the parameters and response strategy library of the fault prediction model are optimized to improve the accuracy of subsequent warnings and responses.
[0025] This invention also includes a dynamic quantification step of fault risk, using a formula. Calculate the fault risk value, where R is the overall fault risk value. The hardware loss weight is P, where P is the actual hardware runtime. For the design lifespan of hardware, The influence of temperature is weighted, where T is the current ambient temperature. For safe temperature threshold, This is the limiting temperature threshold. The load percentage is the weight, and L is the current device load rate. For maximum load rate, Here, W represents the power fluctuation weight, and W represents the actual voltage fluctuation amplitude. Using a standard voltage fluctuation threshold, and through multi-dimensional parameter weighting and fusion, the fault risk under different scenarios is accurately quantified, providing data support for the classification of early warning levels.
[0026] This invention also includes a dynamic early warning threshold calibration step. Based on historical fault data and real-time operating status, the early warning thresholds of each parameter are calibrated periodically. During the calibration process, the distribution pattern of parameters and the correlation with faults are analyzed. For parameters with frequent fluctuations, an adaptive threshold range is adopted, while for stable parameters, a fixed threshold is adopted. At the same time, the threshold range is adjusted in combination with seasonal changes, business peaks and other scenario factors. The calibration results are synchronized to the early warning analysis model in real time to reduce false alarms or missed alarms caused by fixed thresholds.
[0027] This invention also includes a rapid hardware fault location step, employing blockchain-based device identification technology to assign a unique identifier to each hardware component, linking it to component production information, maintenance records, and real-time operation logs to form a full lifecycle data chain. When a fault occurs, the feature vector output by the fault prediction model is matched against a hardware fault feature library to quickly pinpoint suspected faulty components. For clustered server deployments, a distributed fault diagnosis algorithm is used to narrow down the fault range to specific slots or modules through comparison of operational data from adjacent nodes and analysis of temporal characteristics. Thermal imaging-assisted location technology is introduced to capture the temperature distribution of hardware components in real time, identifying localized overheating areas caused by poor contact, short circuits, etc., with location time controlled within seconds. Simultaneously, a hardware health profile is established, predicting vulnerable components based on historical fault data and issuing preventative warnings in advance, buying time for response measures.
[0028] This invention also includes a load intelligent migration optimization step. After a failure occurs, the resource scheduling center scans the remaining computing power, memory capacity, network bandwidth, and storage I / O performance of healthy nodes in the cluster in real time to select target servers that meet the operational requirements of the failed business. Based on virtual machine migration technology and container orchestration protocols, an incremental migration mode is adopted, prioritizing the migration of core business data and session states, while non-core data is asynchronously synchronized subsequently, shortening business interruption time. Migration priorities are set for different business types, with core transaction businesses having the highest priority and adopting a dual-active backup migration mode to achieve zero data loss; non-core computing businesses adopt a load balancing migration mode, distributing them to multiple healthy nodes. During the migration process, network routing policies are dynamically adjusted, traffic shaping technology is enabled to avoid network congestion, and migration progress and target node load changes are monitored in real time. If the target node load exceeds the limit, it automatically switches to a backup node to ensure the stability and reliability of the migration process.
[0029] This invention also includes emergency control steps for the data center environment. For temperature anomalies, a distributed temperature sensor network is used to accurately locate hotspot areas, and a zoned control strategy is employed in conjunction with a precision air conditioning system to adjust the supply air temperature and speed in the hotspot areas. Simultaneously, localized cooling devices are activated to directly cool the high-load server cluster. For scenarios with large-area temperature exceeding limits, unnecessary high-power devices and idle servers are shut down to reduce overall heat generation in the data center. A heat recovery system is activated to transfer heat from high-temperature areas to low-temperature areas, improving energy efficiency. For humidity anomalies, based on humidity distribution data in different areas of the data center, zoned humidification or dehumidification equipment is controlled to adjust humidity to a safe range. During environmental control, temperature and humidity change curves and equipment operating parameters are monitored in real time. A PID control algorithm is used to dynamically optimize control parameters, pushing environmental parameters back to normal range quickly and reducing the probability of secondary failures caused by environmental fluctuations.
[0030] This invention also includes a dynamic priority sorting step, using a formula. Determine response priorities, where P is the response priority coefficient, S is the number of devices affected by the fault, E is the importance coefficient of the fault's impact on business operations, D is the fault recovery difficulty coefficient, and C is the response execution cost coefficient. Higher priority coefficients result in faster response execution. When calculating priorities, a business continuity weight is introduced, with the importance coefficient of core business operations amplified by a multiple of the actual business value. Simultaneously, the response execution cost coefficient is dynamically adjusted based on real-time resource status. When available resources are scarce, the cost of response measures with high resource occupancy is appropriately increased. Establish a response resource conflict coordination mechanism. When multiple faults occur simultaneously, causing resource contention, resources are allocated according to priority coefficients. High-priority faults prioritize core resources, while low-priority faults employ resource reuse or delayed response strategies to reduce response delays caused by resource conflicts and ensure rapid handling of critical faults.
[0031] This invention also includes a power failure emergency response procedure, which monitors mains voltage, frequency, and phase changes in real time. When a mains power interruption or parameters exceeding safe range are detected, a UPS power switch is immediately triggered, with the switchover time controlled within milliseconds to ensure continuous power supply to core equipment. Simultaneously, a diesel generator emergency power supply process is initiated, monitoring generator fuel level, operating status, and output parameters in real time to ensure the generator starts successfully and connects to the power supply system within a specified time. A power supply strategy is implemented based on load priority, prioritizing power supply to critical equipment such as database servers, core switches, and monitoring systems. Non-core equipment can be temporarily powered off or operate at reduced frequency to minimize power consumption. The remaining UPS battery charge and discharge rate are monitored in real time, and the remaining battery life is predicted based on battery health profiles. When the remaining battery life is insufficient, a load shedding sequence is automatically triggered, gradually cutting off unnecessary loads according to priority to ensure continuous power supply for critical business operations. After mains power is restored, a smooth switching technology is used to gradually switch the load from emergency power to mains power, reducing damage to equipment caused by voltage surges.
[0032] In the present invention, there is also a cross-region linkage response step. For large computing power data centers deployed in multiple regions, a high-speed communication link and resource sharing platform between regions are established, and the servers, storage, network, and power resources in each region are virtualized into a global resource pool. When a major failure occurs in a single region and cannot be resolved independently, a resource request is automatically sent to the resource scheduling centers in adjacent regions, specifying resource parameters such as the required computing power, storage, and network bandwidth. The resource sharing platform uses a latency-aware load distribution algorithm based on the resource occupancy situation and the scope of the failure impact in each region to migrate the service load of the failed region to the healthiest region node with the shortest distance and the lowest latency. At the same time, a failure region isolation mechanism is activated to cut off unnecessary network connections between the failed region and the healthy regions to prevent the spread of the failure. During the cross-region migration process, the warning status and resource occupancy in each region are updated synchronously, and a migration session synchronization mechanism is established to ensure the continuity of business cross-region operation and data consistency, achieving load sharing and fault tolerance for the overall business.
[0033] In the present invention, there is also a model self-optimization and iteration step. The fault prediction model and response strategy library are iteratively optimized regularly. Based on the successful cases and failure experiences in the fault handling logs, a reinforcement learning algorithm is used to train the model. The fault warning accuracy rate, response time, service recovery rate, etc. are used as reward indicators, and untimely warning and response failure are used as penalty indicators to continuously adjust the feature weights and decision logic of the model. A fault simulation training library is established, and diverse fault scenarios, including single faults, compound faults, cascading faults, etc., are generated based on historical fault data for offline training and online fine-tuning of the model to improve the model's adaptability to complex faults. The multi-model fusion technology is introduced, and the prediction results of deep learning models, traditional machine learning models, and expert rule models are combined, and a weighted voting mechanism is used to output the final warning result to reduce the false alarm rate and missed alarm rate of a single model. The optimization period can be dynamically adjusted according to the operating load of the data center. The optimization period is shortened during the business peak period and extended during the low peak period to keep the model in an optimal performance state at all times and continuously improve the accuracy of fault warning and the effectiveness of response measures.
[0034] The following further illustrates the specific implementation manners of the present invention through two embodiments: Embodiment 1: Fault Warning Intelligent Response Application for a Single-Region High-Density Large Computing Power Data Center This embodiment is applied to a single-region high-density large computing power data center, which deploys thousands of server clusters, mainly carrying core services such as artificial intelligence training and large-scale data processing. With high equipment density, high power consumption, and concentrated heat output, it is prone to hardware failures, temperature anomalies, network congestion, etc. Fault warning responses are required to be fast and accurate to ensure business continuity. The specific implementation process is as follows: I. Core Processes and Key Steps: Multi-dimensional Data Acquisition and Deployment: Deploy a distributed sensor network in the data center server room. Each server is equipped with a monitoring agent to collect hardware operating parameters in real time, including CPU load, memory usage, hard disk I / O rate, motherboard voltage, etc. Environmental parameters cover temperature, humidity, and airflow speed in various areas of the server room. Power parameters include input voltage, operating current, and power loss. Network parameters include bandwidth utilization, data packet latency, and packet loss rate.
[0035] 2. The data acquisition frequency is dynamically adjusted. The core server acquires data once every 1 second, the ordinary server acquires data once every 5 seconds, and the environmental and power parameters are acquired once every 2 seconds. The data is then transmitted to the data processing node via industrial Ethernet.
[0036] III. Data Preprocessing Execution: A sliding window filtering algorithm with a window size of 10 acquisition cycles is used to denoise the raw data and filter out transient fluctuations. Normalization techniques are used to convert different dimensional parameters to the 0-1 range; CPU load 0-100% corresponds to 0-1, and voltage 200-240V corresponds to 0-1. Anomaly detection thresholds are set using the Isolation Forest algorithm to identify data points deviating from the normal distribution and mark them as suspicious fault features, generating a standardized dataset. Edge nodes perform preliminary preprocessing on local server data, removing obvious outliers before uploading the data to the cloud for centralized processing, achieving edge-cloud collaboration.
[0037] IV. Fault Risk Quantification and Early Warning Analysis: Through formulas Calculate the failure risk value and set the hardware loss weight. =0.3, actual hardware runtime P=36 months, hardware design lifespan =60 months, temperature influence weighting =0.35, current ambient temperature T=32℃, safe temperature threshold =24℃, extreme temperature threshold =40℃, load percentage weight =0.25, current equipment load rate L=90%, maximum load rate =100%, Power fluctuation weight =0.1, actual voltage fluctuation amplitude W=2V, standard voltage fluctuation threshold =1V. The calculated R = 0.78.
[0038] 5. Construct a fault prediction model based on deep learning, input a standardized dataset for feature extraction, and divide the warning level according to the risk value R. R≥0.8 is a level 1 warning, 0.6≤R<0.8 is a level 2 warning, and R<0.6 is a level 3 warning. The dynamic warning threshold is calibrated every 24 hours based on historical fault data.
[0039] VI. Quick Location and Response Strategy Matching for Hardware Failures: Adopt the device identity identification technology based on blockchain to assign unique identifiers to components such as the CPU, memory, and hard disk of each server, and associate component production information, maintenance records, and real-time operation logs. When a failure occurs, the failure prediction model outputs a feature vector, matches the hardware failure feature library, and locks the suspected failed hard disk. By comparing the hard disk IO rate and timing characteristics of adjacent servers through a distributed fault diagnosis algorithm, narrow down the failure range to a specific hard disk slot.
[0040] VII. Introduce thermal imaging technology to capture the temperature distribution in the hard disk area, identify local overheating areas, and the positioning time is controlled within 3 seconds. In the intelligent response strategy matching database, the hardware failure corresponds to load migration and device offline measures, and the core business load migration has the highest priority. Load intelligent migration and cross-device linkage execution: The resource scheduling center continuously scans the remaining computing power, memory capacity, network bandwidth, and storage IO performance of healthy nodes in the cluster, and filters out 3 target servers that meet the business operation requirements. Based on virtual machine migration technology and container orchestration protocol, adopt the incremental migration mode, prioritize the migration of core business data and session states, with the transmission rate controlled at 100MB per second, and non-core data is asynchronously synchronized.
[0041] VIII. Enable traffic shaping technology during the migration process to limit the migration bandwidth occupancy not exceeding 30% of the total bandwidth, and continuously monitor the load changes of target nodes. When the load of a certain target node exceeds 80%, automatically switch to the standby node. Cross-device linkage issues instructions through industrial Ethernet to start the offline process of the failed server, and at the same time联动冷却系统调整对应区域送风参数.
[0042] IX. Emergency Regulation of Computer Room Environment and Verification of Response Effect: The distributed temperature sensing network locates the hot spot area in the middle of the server cluster, and联动精密空调系统将该区域送风温度从24℃调整至22℃,风速提升20%,启动局部冷却装置直接对故障服务器周边降温. Shut down 2 idle servers and 3 non-essential high-power devices to reduce heat generation, and enable the heat recovery system to transfer the heat in the high-temperature area.
[0043] X. Dynamically optimize parameters using the PID control algorithm during the environment regulation process, and adjust the air supply parameters every 10 seconds. Continuously collect relevant parameters after the response is executed. Compare that the CPU load drops from 90% to 65% and the temperature drops from 32℃ to 25℃ before and after the failure, and evaluate that the response measures are effective. Log Record and Model Self-Optimization: Detail the failure occurrence time as 14:30, the failure type as hard disk failure, the warning level as level II, and the response measures include load migration, device offline, and environment regulation. The execution process and processing results are all recorded in the fault handling log.
[0044] A fault prediction model is trained using reinforcement learning algorithms based on log data. The accuracy of early warning and response time are used as reward indicators. The model feature weights are adjusted and the optimization period is set to the off-peak period from 2:00 to 4:00 AM.
[0045] Table 1: Comparison of Fault Handling Performance of High-Density Data Centers in Single Regions
[0046] Table 1 illustrates the advantages of this invention in single-area high-density data centers. Traditional methods rely on manual troubleshooting, resulting in a low early warning accuracy of only 75%, requiring over 5 minutes for fault location, leading to business interruptions of up to 10 minutes, and a high rate of secondary failures. This invention, through multi-dimensional data collection and deep learning models, improves early warning accuracy to over 98%, reduces hardware fault location time to 3 seconds, and enables intelligent load migration and cross-device linkage to keep business interruption time within 30 seconds. Emergency control of the data center environment employs zoned precise cooling and PID optimization, restoring environmental parameters within 5 minutes and reducing the secondary failure rate to below 2%, comprehensively ensuring the continuous operation of core businesses and meeting the stringent requirements of high-density data centers.
[0047] Example 2: Intelligent Response Application for Fault Early Warning in Multi-Regional Distributed High-Computing-Power Data Centers This embodiment is applied to a multi-regional distributed high-computing-power data center, comprising three geographically dispersed regions. Each region deploys server clusters, interconnected via high-speed communication links, to support cross-regional businesses such as financial transactions and government data processing. It is required to handle scenarios such as power outages, cross-regional network interruptions, and overall regional failures to ensure the continuity of cross-regional business operations. The specific implementation process is as follows: I. Core Processes and Key Steps: Multi-dimensional Data Acquisition and Cross-Regional Synchronization: Distributed sensor networks and server monitoring agents are deployed in data centers in each region to collect hardware operating parameters, environmental parameters, power parameters, and network load data, with the acquisition frequency consistent with that of a single region. Real-time data synchronization across regions is achieved through high-speed encrypted communication links, with synchronization latency controlled within 10 milliseconds. A global data sharing platform is established to centrally store monitoring data from each region. Data preprocessing nodes are set up in each region to perform noise reduction, normalization, and anomaly identification locally, generating standardized datasets which are then uploaded to the global data processing center. Dynamic Early Warning Threshold Calibration and Early Warning Analysis: Based on historical fault data and real-time operating status, early warning thresholds for each parameter are calibrated every 24 hours.
[0048] For parameters with frequent fluctuations, such as network latency, an adaptive threshold range is used, with the range width dynamically adjusted based on the fluctuation amplitude over the past 7 days. For stable parameters, such as environmental humidity, a fixed threshold is used. A global fault prediction model is constructed, inputting standardized datasets from each region to extract cross-regional fault correlation features and identify the risk of a fault in a single region spreading to other regions.
[0049] Through formula Determine the response priority, set the number of devices affected by the fault to S=200, the importance coefficient of the fault impacting core financial business to E=5, the fault recovery difficulty coefficient to D=3, and the response execution cost coefficient to C=2. The calculated P==166.67, indicating a high priority coefficient, thus the response measures should be executed first.
[0050] Power Failure Emergency Response and Cross-Equipment Interoperability: Real-time monitoring of mains voltage, frequency, and phase changes; when a mains power outage is detected in a region, UPS power switching is immediately triggered, with the switching time controlled within 5 milliseconds to ensure continuous power supply to core equipment. Simultaneously, the diesel generator emergency power supply process is initiated, with real-time monitoring of generator fuel level, operating status, and output parameters to ensure startup and connection to the power supply system within 30 seconds.
[0051] Power supply is prioritized based on load priority, ensuring power to database servers, core switches, and monitoring systems while cutting off power to 10 non-core computing devices to reduce power consumption. The system monitors the remaining battery power and discharge rate of the UPS in real time, predicts battery life based on battery health profiles, and triggers a load shedding sequence when battery life is less than 30 minutes, gradually cutting off unnecessary loads according to priority.
[0052] Cross-regional coordinated response and load migration: When a mains power outage occurs in Region 1 and cannot be resolved independently, the global resource sharing platform automatically sends resource requests to Regions 2 and 3, specifying the required computing power, storage, and network bandwidth parameters.
[0053] The resource sharing platform uses a latency-aware load distribution algorithm to migrate the core financial business load of Region 1 to the nearest and lowest latency healthy node in Region 2, while non-core business is migrated to Region 3.
[0054] Initiate a fault area isolation mechanism to sever unnecessary network connections between Region 1 and other regions to prevent the fault from spreading. During cross-region migration, establish a migration session synchronization mechanism, employ encrypted transmission to ensure data security, and update the early warning status and resource usage of each region in real time to ensure the continuity of business operations and data consistency across regions.
[0055] Rapid hardware fault location and environmental control: A CPU fault warning was issued for a server in Region 2. Using blockchain-based device identification technology, CPU production information, maintenance records and real-time operation logs were linked. Combined with the feature vector output by the fault prediction model, the hardware fault feature library was matched to locate the faulty CPU.
[0056] By comparing the operational data and timing characteristics of adjacent nodes using a distributed fault diagnosis algorithm, the fault range was narrowed down to the CPU socket. Thermal imaging technology was introduced to capture the temperature distribution in the CPU area, identifying localized overheating areas with a location time controlled within 4 seconds. For the abnormal temperature in the second server room, a zoned control strategy was implemented in conjunction with the precision air conditioning system. This involved adjusting the supply air temperature and speed in hotspot areas, activating local cooling devices, shutting down three idle servers, activating the heat recovery system, and using a PID control algorithm to dynamically optimize control parameters, allowing environmental parameters to quickly return to normal ranges.
[0057] Response effect verification and log recording: After the response is executed, parameters of each area are continuously collected. After the power is restored in area 1, the load is switched from emergency power to mains power using smooth switching technology. The business load corresponding to the faulty CPU in area 2 has been migrated and is running stably.
[0058] Comparing the data before and after the fault, core services in Region 1 remained uninterrupted, CPU load in Region 2 returned to normal, and temperatures returned to safe ranges, indicating the effectiveness of the response measures. A detailed record was kept of the fault's occurrence time, type, warning level, response measures, execution process, and results, creating a global fault handling log, which was synchronized to all regional data centers.
[0059] Model self-optimization iteration: Based on successful cases and failure experiences in the global fault handling log, a reinforcement learning algorithm is used to train the global fault prediction model. The early warning accuracy, response time, and business recovery rate are used as reward indicators, while failure to issue timely warnings and failure to respond are used as penalty indicators. The model feature weights and decision logic are adjusted accordingly.
[0060] A fault simulation training library is established to generate diverse fault scenarios based on historical fault data, including single faults, compound faults, and cascading faults. The model is then trained offline and fine-tuned online. Multi-model fusion technology is introduced, combining the prediction results of deep learning models, traditional machine learning models, and expert rule models. A weighted voting mechanism is used to output the final early warning result. The optimization cycle is dynamically adjusted according to the data center's operating load: every 12 hours during peak business periods and every 24 hours during off-peak periods.
[0061] Table 2: Comparison of Fault Handling Performance of Multi-Region Distributed Data Centers
[0062] Table 2 shows that traditional methods for handling cross-regional fault responses rely on manual coordination, resulting in response times as long as 30 minutes, a power outage service continuity rate of only 70%, a cross-regional load migration success rate as low as 80%, and a high probability of fault propagation. This invention, through a global resource sharing platform and a cross-regional linkage mechanism, shortens the cross-regional fault response time to 5 minutes, achieving a power outage emergency response service continuity rate of over 99%.
[0063] Cross-region load migration employs latency-aware algorithms and session synchronization mechanisms, increasing the success rate to 99.5%, while fault area isolation mechanisms reduce the probability of propagation to below 1%. Global model self-optimization and multi-model fusion technologies reduce the false alarm rate to below 3%, comprehensively ensuring the continuity of core business operations across regions and adapting to the complex needs of multi-regional distributed data centers.
[0064] Reference Figure 2 This diagram visually demonstrates the high accuracy of this invention's early warning system across various fault scenarios, stemming from the innovative fusion of multi-dimensional data acquisition and deep learning models. Traditional methods rely on single-parameter monitoring and simple threshold judgments, failing to adequately capture the characteristics of complex faults. Cross-regional fault warning accuracy is only 65%, insufficient to meet the needs of multiple scenarios. This invention utilizes a distributed sensor network to cover hardware, environment, power, and network parameters across all dimensions. Combined with the Isolation Forest algorithm to eliminate data noise, and a deep learning model to accurately extract fault features, further optimized through dynamic threshold calibration, this ensures that the accuracy of early warning for various fault scenarios remains above 95%. Especially in complex scenarios such as cross-regional faults, the accuracy still reaches 95% thanks to cross-regional data synchronization and correlation feature analysis, significantly reducing the risk of false alarms and missed alarms, laying the foundation for rapid response.
[0065] Reference Figure 3 This diagram highlights the high reliability advantage of this invention's cross-regional linkage, which hinges on the innovation of high-speed communication links and intelligent load balancing algorithms. Traditional methods are limited by communication latency and unreasonable resource scheduling; as the migration distance increases, the success rate drops significantly, reaching only 60% at 100km, failing to guarantee cross-regional business continuity. This invention establishes a high-speed encrypted communication link, controlling cross-regional data synchronization latency to within 10 milliseconds; the global resource sharing platform employs a latency-aware load balancing algorithm, prioritizing healthy nodes with close proximity and low latency; and session synchronization and encrypted transmission mechanisms are enabled during migration to ensure data consistency and security. Even at a migration distance of 100km, the success rate remains above 95%, effectively solving the challenge of cross-regional fault handling in multi-regional distributed data centers.
[0066] Reference Figure 4This diagram illustrates the stable response capability of this invention under high-load scenarios, with the key being edge computing and resource scheduling optimization. Traditional methods rely on centralized cloud processing. As business load increases, server computing power becomes strained, and response latency increases dramatically, reaching 500 milliseconds at 95% high load, failing to meet the low-latency requirements of core businesses. This invention supports local preprocessing of critical data at edge nodes, reducing cloud pressure; the resource scheduling center scans the remaining computing power of healthy nodes in real time and dynamically allocates response resources; and traffic shaping technology is used during load migration to avoid network congestion. Even under 95% high load scenarios, the response latency is still controlled within 20 milliseconds, with minimal latency fluctuations across different load rates, ensuring rapid fault response during high-load operation of large-scale computing data centers, and adapting to the stringent requirements of core businesses such as artificial intelligence training and financial transactions.
[0067] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart response method for fault early warning in high-computing-power data centers, characterized in that, Includes the following steps: The data acquisition process involves deploying a distributed sensor network and server monitoring agents to collect hardware device operating parameters, data center environment parameters, power supply parameters, and network load data in real time. The data acquisition frequency is dynamically adjusted based on the importance of the equipment. In the preprocessing step, a sliding window filtering algorithm is used to reduce noise in the collected data. Normalization technology is used to unify the numerical range of parameters with different dimensions. The isolated forest algorithm is combined to identify outliers in the data, mark suspicious fault features, and generate a standardized dataset. The early warning analysis steps involve constructing a fault prediction model based on deep learning, inputting a standardized dataset for feature extraction and fault type matching, classifying early warning levels according to the probability of fault occurrence, scope of impact, and difficulty of recovery, and setting early warning thresholds. The strategy matching step involves establishing a mapping database between fault types and response plans, and matching corresponding response measures for different fault types. Cross-device linkage execution steps: Through the industrial Ethernet and device control interface, response commands are sent to the devices to realize the coordinated action of multiple devices, and the execution status is fed back in real time during the linkage process; The response effectiveness verification step involves continuously collecting relevant parameters after the response is executed, comparing the parameter changes before and after the fault is resolved, evaluating the effectiveness of the response measures, and triggering a secondary response if the fault is not eliminated, adjusting the response strategy or escalating the processing level. The logging and parameter update process records the fault occurrence time, fault type, warning level, response measures, execution process and processing results in detail, forming a fault handling log. Based on the log data, the parameters and response strategy library of the fault prediction model are optimized.
2. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a dynamic quantification step for fault risk, using formulas. Calculate the fault risk value, where R is the overall fault risk value. The hardware loss weight is P, where P is the actual hardware runtime. For the lifespan of hardware design, The influence of temperature is weighted, where T is the current ambient temperature. For safe temperature threshold, This is the limiting temperature threshold. The load percentage is the weight, and L is the current device load rate. For maximum load rate, Here, W represents the power fluctuation weight, and W represents the actual voltage fluctuation amplitude. This is the standard voltage fluctuation threshold.
3. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a dynamic early warning threshold calibration step, which periodically calibrates the early warning thresholds of each parameter based on historical fault data and real-time operating status. During the calibration process, the distribution pattern of parameters and the correlation with faults are analyzed. For parameters with frequent fluctuations, an adaptive threshold range is used, while for stable parameters, a fixed threshold is used. At the same time, the threshold range is adjusted in combination with seasonal changes and peak business scenarios. The calibration results are synchronized to the early warning analysis model in real time.
4. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a rapid hardware fault location step, which uses blockchain-based device identity identification technology to assign a unique identifier to each hardware component and associate it with component production information, maintenance records and real-time operation logs to form a full lifecycle data chain. When a fault occurs, the feature vector output by the fault prediction model is combined with the hardware fault feature library to quickly locate the suspected faulty component. For clustered servers, a distributed fault diagnosis algorithm is used to narrow down the fault range to a specific slot or module by comparing the running data of adjacent nodes and analyzing the time series features.
5. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a load intelligent migration optimization step. After a failure occurs, the resource scheduling center scans the remaining computing power, memory capacity, network bandwidth and storage IO performance of healthy nodes in the cluster in real time, selects target servers that meet the requirements of the failed business operation, and adopts an incremental migration mode based on virtual machine migration technology and container orchestration protocol to prioritize the migration of core business data and session state.
6. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes emergency control procedures for the data center environment. In response to abnormal temperature faults, a distributed temperature sensor network is used to accurately locate hot spots and link the precision air conditioning system to adopt a zone control strategy to adjust the air supply temperature and wind speed corresponding to the hot spots. At the same time, local cooling devices are activated to directly cool down the high-load server cluster.
7. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a dynamic priority ranking step, using a formula. Determine the response priority, where P is the response priority coefficient, S is the number of devices affected by the fault, E is the importance coefficient of the fault's impact on business, D is the fault recovery difficulty coefficient, and C is the response execution cost coefficient.
8. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes power failure emergency response steps, real-time monitoring of mains voltage, frequency and phase changes, and immediate triggering of UPS power switching when a mains power interruption or parameters exceeding the safe range are detected, with switching time controlled in milliseconds; at the same time, the diesel generator emergency power supply process is started, and the generator fuel quantity, operating status and output parameters are monitored in real time. Based on load priority, the power supply strategy prioritizes the power supply to critical equipment such as database servers, switches, and monitoring systems. It monitors the remaining power and discharge rate of UPS batteries in real time, predicts the battery life based on battery health profiles, and automatically triggers a load shedding sequence when the battery life is insufficient, gradually cutting off unnecessary loads according to priority.
9. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes cross-regional linkage response steps. For high-performance data centers deployed in multiple regions, a high-speed communication link and resource sharing platform are established between regions, and the servers, storage, network and power resources of each region are virtualized into a global resource pool. When a major failure occurs in a single region and cannot be resolved independently, a resource request is automatically sent to the resource scheduling center of the adjacent region, specifying the required computing power, storage, and network bandwidth, etc. Based on the resource usage and fault impact range of each region, the resource sharing platform uses a latency-aware load distribution algorithm to migrate the business load of the faulty region to the nearest healthy region node with the lowest latency.
10. The intelligent response method for fault early warning in high-computing-power data centers according to claim 1, characterized in that, It also includes a model self-optimization iteration step, regularly iterating and optimizing the fault prediction model and response strategy library, and using reinforcement learning algorithms to train the model based on successful cases and failure experiences in the fault handling log; The model's feature weights and decision logic are continuously adjusted by using fault warning accuracy, response time, and business recovery rate as reward indicators, and failure to issue timely warnings or fail to respond as penalty indicators. Establish a fault simulation training library, generate diverse fault scenarios based on historical fault data, including single faults, compound faults, and cascading faults, and perform offline training and online fine-tuning of the model.