Data center operation and maintenance management method and management system based on artificial intelligence
Through dynamic learning of protocol features and timing optimization based on artificial intelligence, the power load conduction path is predicted, and the systemic failure problem in the upgrade of heterogeneous equipment in data centers is solved, intelligent operation and maintenance management is realized, and operation and maintenance efficiency and system stability are improved.
Patent Information
- Application Number
- CN202510448919.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-10
AI Technical Summary
In the data center, during the firmware upgrade of heterogeneous devices, due to the timing offset and power strategy failure caused by the differences in the implementation of the private expansion of the Redfish protocol by BMCs of various manufacturers, the complexity and systemic risks of operation and maintenance are increased. It is difficult for the existing technology to dynamically learn protocol extension characteristics and predict power load conduction paths, resulting in insufficient operation and maintenance automation level and reliability.
The protocol feature dynamic learning module, timing orchestration optimization module, cascade failure protection module and multimodal abnormality self-healing module are adopted to dynamically analyze the BMC protocol characteristics through reinforcement learning and deep learning technology, optimize the instruction sequence, predict the power consumption conduction path, generate load shaping strategies, and perform self-healing operations in abnormal situations to achieve intelligent management.
It significantly improves the automation level and system stability of data center operation and maintenance, reduces the operation and maintenance cycle, avoids hidden risks in the coexistence scenario of cross-brand equipment, and ensures the reliability and efficiency of the system.
Smart Images

Figure CN120295881A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data center operation and maintenance management, and specifically relates to an artificial intelligence-based data center operation and maintenance management method and management system. Background Art
[0002] With the large-scale development of data centers, modern data centers generally adopt a hybrid architecture of multi-brand servers to achieve a balance between cost and performance. However, in key operation and maintenance scenarios such as firmware upgrades, the differences in the private extensions of the Redfish protocol implemented by each vendor's baseboard management controller (BMC) have become a major technical bottleneck restricting operation and maintenance automation. When implementing the Redfish protocol, mainstream vendors (such as HP iLO, Dell iDRAC, and Inspur BMC) generally introduce private instruction sets and response timing logics. For example, iLO v2.3 requires inserting a power policy negotiation window of 300±50ms before sending the FirmwareUpdate instruction, while iDRAC v4.2 forcibly requires maintaining continuous heartbeat packets (interval ≤ 2 seconds) during the firmware transmission stage. This millisecond-level timing difference causes uncontrollable offsets in the state machines of different brand devices in the same batch of upgrade operations, significantly increasing the complexity of operation and maintenance.
[0003] In addition, when the upgrade timings of heterogeneous devices are misaligned beyond the synchronization fault tolerance window of the power distribution unit (PDU) (usually 200 - 500ms), it will trigger an avalanche failure of the power policy. Typical cases show that when performing a full-scale upgrade in a hybrid cabinet (including 10 units of HP / Dell / Inspur devices each), the instantaneous power consumption peak of iLO devices (about 300W / unit) and the firmware verification cycle of iDRAC (concentrated power consumption of 150W / unit every 5 seconds) may produce a superimposed effect, causing the PDU to trigger cascading overload protection within 3.2 seconds, resulting in accidental power-off of the entire row of cabinets. This conduction mechanism of cascading failure further exacerbates the systemic risk.
[0004] Existing artificial intelligence technologies have obvious limitations in solving the above problems. Although traditional rule-based automated operation and maintenance systems can identify obvious protocol errors (such as abnormal HTTP status codes), they cannot dynamically learn protocol extension features and still rely on manual reverse engineering to maintain the vendor instruction set knowledge base. Facing new BMC firmware (such as iLO 6 with AI acceleration function), the feature extraction period may be as long as several months, which is difficult to meet the requirements of rapid iteration. At the same time, existing machine learning models (such as LSTM networks) are mostly used for macro resource scheduling and lack the ability to model microsecond-level instruction interaction timings, making it difficult to predict the cascading effects of cross-vendor operations.
[0005] Industry mainstream solutions usually adopt manual batch operations (such as dividing upgrade batches by brand), but this not only leads to a more than three-fold extension of the operation and maintenance cycle but also fails to avoid the hidden risks in the scenario of coexistence of cross-brand devices. Therefore, there is an urgent need for an artificial intelligence-based data center operation and maintenance method that fundamentally solves the systematic failure problem caused by firmware upgrades of heterogeneous devices and improves the automation level and reliability of data center operation and maintenance by deeply analyzing protocol extension features, dynamically predicting the power load conduction path, and intelligently orchestrating the microsecond-level operation timing. Summary of the Invention
[0006] The purpose of the present invention is to provide an artificial intelligence-based data center operation and maintenance management method, which solves the systematic failure problem of heterogeneous devices in key operation and maintenance scenarios such as firmware upgrades through intelligent instruction coordination, risk prediction, and anomaly self-healing mechanisms.
[0007] To achieve the above purpose, the present invention provides the following technical solution: An artificial intelligence-based data center operation and maintenance management method, the operation and maintenance management method includes a protocol feature dynamic learning module, a timing orchestration optimization module, a cascading failure protection module, and a multimodal anomaly self-healing module;
[0008] The protocol feature dynamic learning module is used to collect the Redfish protocol extension response timing characteristics of multi-vendor baseboard management controllers (BMCs) and build a dynamically updated protocol knowledge base;
[0009] The timing orchestration optimization module is used to generate an optimal instruction emission sequence according to the protocol knowledge base and insert a dynamic buffer interval in combination with real-time power distribution unit (PDU) load data;
[0010] The cascading failure protection module is used to predict the microsecond-level power consumption conduction path through a spatio-temporal graph attention network (ST-GAT) and generate a load shaping strategy to avoid overload risks;
[0011] The multimodal anomaly self-healing module is used to synchronously analyze BMC log streams, PDU telemetry data, and network packet characteristics, and generate an incremental rollback plan using a decision forest driven by reinforcement learning.
[0012] A reinforcement learning agent is set in the protocol feature dynamic learning module. The reinforcement learning agent interacts with multi-vendor BMCs. When a new protocol extension feature is detected, the reinforcement learning agent automatically extracts its timing logic and updates the protocol knowledge base, and at the same time generates a dynamic buffer time recommendation and transmits it to the timing orchestration optimization module.
[0013] The timing scheduling optimization module includes a millisecond-level convolutional branch, a second-level convolutional branch, and a quantum annealing solver. The millisecond-level convolutional branch is used to extract short-term sensitive operation features, the second-level convolutional branch is used to capture long-term dependencies, and the quantum annealing solver solves the optimal instruction sequence that minimizes the variance of PDU load fluctuations by mapping the device operation timing combination optimization problem. The specific formula is as follows:
[0014]
[0015] Where P i (t) represents the instantaneous power consumption of the i-th device at time t, represents the average power consumption of all devices at time t, N represents the total number of devices, and T represents the total duration. Through this formula, precise control of PDU load fluctuations is achieved, ensuring the stability of device operation timing.
[0016] A spatio-temporal graph attention network (ST-GAT) is set in the cascading failure protection module. The time sliding window is set to 1 second, and the spatial adjacency matrix is constructed based on the circuit topology distance. The multi-head attention mechanism is used to capture the microsecond-level power consumption conduction path across cabinets. When an overload risk is predicted, the gradient backpropagation algorithm is used to calculate the device operation delay. The specific formula is as follows:
[0017]
[0018] Where Δt i represents the operation delay of the i-th device, represents the maximum allowable load of the power distribution unit, P total (t) represents the total load at time t, represents the power consumption change rate of the i-th device. Through this formula, the device operation timing is dynamically adjusted to avoid triggering the protection mechanism due to overload.
[0019] A fusion Transformer architecture is set in the multi-modal anomaly self-healing module. The Transformer is used to synchronously analyze the BMC log stream, PDU telemetry data, and network packet characteristics. When a cross-layer anomaly pattern is detected, an incremental rollback scheme is generated through a decision forest driven by reinforcement learning. The splitting threshold of each tree node in the decision forest is dynamically optimized through Q-Learning. The specific formula is as follows:
[0020]
[0021] Where Q(s, a) represents the expected reward for taking action a in the current state s, α represents the learning rate, r represents the immediate reward, γ represents the discount factor, and s′ represents the next state. Through this formula, the rollback strategy is optimized with the goal of minimizing the scope of affected devices.
[0022] In addition, the present invention also provides an operation and maintenance management system for a data center based on artificial intelligence, and the intelligent operation and maintenance management system is applicable to the method for operation and maintenance management of a data center based on artificial intelligence. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic diagram of the module structure of the method for operation and maintenance management of a data center based on artificial intelligence in an embodiment of the present invention.
[0024] Figure 2 It is a flowchart of the operation of the timing arrangement optimization module in an embodiment of the present invention.
[0025] Figure 3 It is a schematic diagram of the structure of the spatio-temporal graph attention network model of the cascading failure protection module in an embodiment of the present invention.
[0026] Figure 4 It is a working principle diagram of the fusion Transformer architecture and the reinforcement learning-driven decision forest of the multi-modal anomaly self-healing module in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] The present invention provides a method for operation and maintenance management of a data center based on artificial intelligence. The core lies in the collaborative work of a protocol feature dynamic learning module, a timing arrangement optimization module, a cascading failure protection module, and a multi-modal anomaly self-healing module to achieve intelligent management of heterogeneous devices in a data center in key operation and maintenance scenarios. The following will elaborate on the specific implementation manners of the present invention in detail in conjunction with the drawings and specific embodiments.
[0029] Such as Figure 1As shown in the figure, the module structure of the present invention includes a protocol feature dynamic learning module, a timing arrangement optimization module, a cascading failure protection module, and a multi-modal anomaly self-healing module. The protocol feature dynamic learning module transmits a dynamic buffer time recommendation value to the timing arrangement optimization module. After the timing arrangement optimization module generates an instruction sequence, the cascading failure protection module provides real-time feedback on the load shaping strategy, and the multi-modal anomaly self-healing module generates a rollback instruction according to the global state, forming a closed-loop control flow. The specific implementation method starts with the protocol feature dynamic learning module. The core function of this module is to collect the Redfish protocol extension response timing characteristics of multi-vendor baseboard management controllers (BMCs) and build a dynamically updated protocol knowledge base. In practical applications, a data center usually contains devices from different vendors. The BMCs of these devices support the Redfish protocol, but each vendor may have extended the protocol to varying degrees. To address this heterogeneity, this module sets up a reinforcement learning agent. This agent automatically extracts new protocol extension features and updates the protocol knowledge base by interacting with multi-vendor BMCs. For example, when a certain vendor adds a specific firmware upgrade instruction sequence, the reinforcement learning agent can detect this change and incorporate it into the knowledge base by analyzing its timing logic. The protocol feature dynamic learning module stores the Redfish protocol extension response timing characteristics of multi-vendor baseboard management controllers (BMCs) using a graph database (Neo4j). At the same time, this module also generates a dynamic buffer time recommendation based on the new features for use by subsequent modules. The formula for calculating the recommended value of the dynamic buffer time is as follows:
[0030]
[0031] where Δt i represents the average response delay of the i-th protocol extension feature, and M represents the number of current protocol extension features. Through this formula, the rationality of the dynamic buffer time can be ensured, thus providing a stable timing basis for subsequent operations.
[0032] Next, the timing arrangement optimization module generates an optimal instruction emission sequence based on the protocol knowledge base generated by the protocol feature dynamic learning module and inserts a dynamic buffer interval in combination with the real-time power distribution unit (PDU) load data. As Figure 2 shown, this module includes a millisecond-level convolution branch, a second-level convolution branch, and a quantum annealing solver. The millisecond-level convolution branch is used to extract short-term sensitive operation characteristics, such as operations that require quick response during device startup or firmware upgrade; the second-level convolution branch is used to capture long-term dependencies, such as performance fluctuations that may occur during long-term device operation. The quantum annealing solver is responsible for mapping the device operation timing combination optimization problem into a mathematical model and ensuring system stability by solving the optimal instruction sequence that minimizes the variance of PDU load fluctuations. The specific optimization formula is as follows:
[0033]
[0034] Among them, P i (t) represents the instantaneous power consumption of the i-th device at time t, represents the average power consumption of all devices at time t, N represents the total number of devices, and T represents the total duration. Through this formula, the PDU load fluctuation can be precisely controlled, avoiding system instability caused by uneven load. In practical applications, assume that a data center has 100 servers, and the power consumption range of each server is from 200W to 500W. The optimal instruction sequence generated by the quantum annealing solver can ensure a reasonable distribution of the operation timing of each device, thereby controlling the PDU load fluctuation within 5%.
[0035] The cascading failure protection module predicts the microsecond-level power consumption conduction path through a spatio-temporal graph attention network (ST-GAT) and generates a load shaping strategy to avoid overload risks. As Figure 3 shown, the inputs of this module include device firmware features and circuit topologies. The time sliding window of the spatio-temporal graph attention network is set to 1 second, and the spatial adjacency matrix is constructed based on the circuit topology distance. The multi-head attention mechanism is used to capture the power consumption conduction path across cabinets. For example, in a certain data center, if the devices in a cabinet start simultaneously, it may cause a sudden increase in the PDU instantaneous load. At this time, the ST-GAT model will predict this overload risk and calculate the device operation delay amount through the gradient backpropagation algorithm. The specific formula is as follows:
[0036]
[0037] Among them, Δt i represents the operation delay amount of the i-th device, represents the maximum allowable load of the PDU, P total (t) represents the total load at time t, represents the power consumption change rate of the i-th device. Through this formula, the module can dynamically adjust the device operation timing to avoid the overload trigger protection mechanism. For example, when the total load of the devices in a cabinet approaches the maximum allowable load of the PDU, the module will automatically delay the operation of some devices, thereby controlling the load within a safe range.
[0038] The multimodal anomaly self-healing module synchronously analyzes BMC log streams, PDU telemetry data, and network packet characteristics through integrating the Transformer architecture, and uses a reinforcement learning-driven decision forest to generate an incremental rollback plan. The decision forest consists of 100 CART trees, and the node splitting threshold is dynamically optimized through offline pre-training (historical fault dataset) and online fine-tuning (real-time Q-Learning). The state space s is defined as device health (0 - 1), PDU load deviation rate (%), and network packet loss rate (%). The action space a includes rollback priority (high / medium / low) and rollback granularity (single device / rack / cluster). As Figure 4 shown, the core of this module lies in extracting anomaly patterns from multimodal data through the Transformer architecture and optimizing the rollback strategy in combination with reinforcement learning. For example, when the BMC log of a certain device shows an abnormal restart record, while the PDU telemetry data shows a sudden drop in its power consumption, and the network packet characteristics show a communication interruption, the module will determine that the device may have a serious fault. At this time, the reinforcement learning-driven decision forest will generate an incremental rollback plan, and the specific formula is as follows:
[0039]
[0040] Among them, Q(s, a) represents the expected return of taking action a in the current state s, α represents the learning rate, r represents the immediate reward, γ represents the discount factor, and s′ represents the next state. Through this formula, the module can optimize the rollback strategy with the goal of minimizing the scope of affected devices. For example, during a firmware upgrade process, if it is found that some devices cannot run properly due to compatibility issues, the module will give priority to rolling back the firmware versions of the affected devices while retaining the upgrade results of other devices, thereby minimizing the impact on the overall system.
[0041] In addition, the present invention also provides an artificial intelligence-based data center operation and maintenance management system, which is applicable to the above operation and maintenance management method. In actual deployment, this system can support large-scale data centers through a distributed architecture. For example, in a certain ultra-large-scale data center, the system can deploy the protocol feature dynamic learning module on edge nodes close to the device side to reduce data transmission latency; deploy the timing orchestration optimization module and the cascading failure protection module on the central control node to centrally process global optimization tasks; deploy the multimodal anomaly self-healing module in the cloud to make full use of the high-performance computing power of cloud computing. Through this hierarchical deployment method, the system can achieve comprehensive intelligent management of the data center while ensuring real-time performance.
[0042] In summary, through the collaborative work of the protocol feature dynamic learning module, the timing arrangement optimization module, the cascading failure protection module, and the multi-modal anomaly self-healing module, the present invention solves the systematic failure problem of heterogeneous devices in key operation and maintenance scenarios such as firmware upgrade. Through the verification of actual application scenarios, the present invention can significantly improve the operation and maintenance efficiency and system stability of the data center, providing reliable technical support for the intelligent management of large-scale data centers.
[0043] The preferred embodiments of the present invention disclosed above are only used to help illustrate the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to only the specific embodiments. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principle and practical application of the present invention, so that those skilled in the relevant technical field can understand and utilize the present invention well.
Claims
1. A method for operation and maintenance management of a data center based on artificial intelligence, characterized in that: The operation and maintenance management method includes a protocol feature dynamic learning module, a timing arrangement optimization module, a cascading failure protection module, and a multimodal anomaly self-healing module; The protocol feature dynamic learning module is used to collect the Redfish protocol extension response timing features of multi-vendor baseboard management controllers and construct a dynamically updated protocol knowledge base; The timing arrangement optimization module is used to generate an optimal instruction emission sequence according to the protocol knowledge base and insert a dynamic buffer interval in combination with the real-time power distribution unit load data; The cascading failure protection module is used to predict the microsecond-level power consumption conduction path through a spatio-temporal graph attention network and generate a load shaping strategy to avoid overload risks; The multimodal anomaly self-healing module is used to synchronously analyze the baseboard management controller log stream, the power distribution unit telemetry data, and the network packet characteristics, and generate an incremental rollback scheme using a decision forest driven by reinforcement learning.
2. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 1, wherein: A reinforcement learning agent is set in the protocol feature dynamic learning module. The reinforcement learning agent interacts with multi-vendor baseboard management controllers. When a new protocol extension feature is detected, the reinforcement learning agent automatically extracts its timing logic and updates the protocol knowledge base, and at the same time generates a dynamic buffer time recommendation and transmits it to the timing arrangement optimization module.
3. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 1, characterized in that: The timing arrangement optimization module includes a millisecond-level convolution branch, a second-level convolution branch, and a quantum annealing solver; the millisecond-level convolution branch is used to extract short-term sensitive operation features, the second-level convolution branch is used to capture long-term dependencies, and the quantum annealing solver solves the optimal instruction sequence that minimizes the variance of the power distribution unit load fluctuation by mapping the device operation timing combination optimization problem.
4. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 3, characterized in that: The quantum annealing solver solves the optimal instruction sequence that minimizes the variance of the power distribution unit load fluctuation through the following formula: Among them, P i (t) represents the instantaneous power consumption of the i-th device at time t, represents the average power consumption of all devices at time t, N represents the total number of devices, and T represents the total duration.
5. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 1, wherein: A spatio-temporal graph attention network is set in the cascading failure protection module. The spatio-temporal graph attention network inputs the device firmware features and the circuit topology structure. The time sliding window of the spatio-temporal graph attention network is set to 1 second, and the spatial adjacency matrix is constructed based on the circuit topology distance. The cross-cabinet power consumption conduction path is captured through a multi-head attention mechanism. When an overload risk is predicted, the gradient backpropagation algorithm is used to calculate the device operation delay.
6. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 5, characterized in that: The cascading failure protection module calculates the device operation delay through the following formula: Among them, Δt i represents the operation delay of the i-th device, represents the maximum allowable load of the power distribution unit, P total (t) represents the total load at time t, represents the power consumption change rate of the i-th device.
7. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 1, wherein: A fusion Transformer architecture is set in the multimodal anomaly self-healing module. The Transformer is used to synchronously analyze the baseboard management controller log stream, the power distribution unit telemetry data, and the network packet characteristics. When a cross-layer anomaly pattern is detected, an incremental rollback scheme is generated through a decision forest driven by reinforcement learning.
8. The method for operation and maintenance management of a data center based on artificial intelligence according to claim 7, wherein: The decision forest driven by reinforcement learning optimizes the rollback strategy through the following formula: Among them, Q(s,a) represents the expected return of taking action a in the current state s, α represents the learning rate, r represents the immediate reward, γ represents the discount factor, and s′ represents the next state.
9. An operation and maintenance management system for a data center based on artificial intelligence, characterized in that: This intelligent operation and maintenance management system is applicable to the data center operation and maintenance management method based on artificial intelligence described in any one of claims 1 to 8.
10. The data center operation and maintenance management system based on artificial intelligence according to claim 9, characterized in that: The intelligent operation and maintenance management system supports large-scale data centers through a distributed architecture, deploys the protocol feature dynamic learning module on edge nodes close to the device side, deploys the timing orchestration optimization module and the cascading failure protection module on the central control node, and deploys the multi-modal anomaly self-healing module in the cloud.
Citation Information
Patent Citations
Machine learning automatic process management and optimization system and method based on micro-service
CN111913715A
Operation and maintenance management method and system based on artificial intelligence
CN117390536A
Wireless OTA upgrading method and system based on embedded operating system
CN117640623A
Distributed DTU power distribution terminal for Internet of Things system
CN118646772A
Industrial environment information wireless monitoring system based on Internet of Things and control method thereof
CN119324937A
Cited By
Operation and maintenance management method and system of data center
CN121585572A