Automatic operation and maintenance execution system for high-computing-power data center
By integrating multi-source data and constructing a two-dimensional interlocked decision module, the system achieves full-process perception and real-time control of the operation and maintenance system of high-computing-power data centers, solving the problems of data silos and insufficient risk perception, and improving the security and efficiency of operation and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-13
- Publication Date
- 2026-04-10
AI Technical Summary
Existing data center operation and maintenance systems suffer from problems such as data silos, insufficient risk perception and control capabilities, and architectural limitations, making it difficult to cope with the complexity and dynamism of high-performance data centers, resulting in low operation and maintenance efficiency and insufficient security.
The system integrates multi-source data using a data acquisition module, performs precise quantification through an operations and maintenance personnel capability assessment module and a task risk assessment module, constructs a two-dimensional interlocking decision-making module for real-time linkage and control, and combines a closed-loop update module and model iterative updates to form a full-process, dynamically interconnected operations and maintenance system.
It achieves comprehensive perception of "human-machine-environment-task", accurate assessment and scientific decision-making, dynamically adapts to operation and maintenance needs, improves the security, accuracy and efficiency of operation and maintenance, and adapts to the complex scenarios and dynamic changes of data centers.
Smart Images

Figure CN121836306A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data center operation and maintenance control technology, and in particular to an automated execution system for the operation and maintenance of high-computing-power data centers. Background Technology
[0002] High-performance data centers, as the core infrastructure of the digital economy, bear critical tasks such as massive data storage and high-performance computing. Their operational quality directly determines business continuity and data security. With the exponential growth in computing power demand, data center scale continues to expand, equipment density increases, and operational scenarios become increasingly complex, encompassing multiple dimensions such as environmental monitoring, equipment maintenance, personnel management, and task scheduling. This places extremely high demands on the accuracy, real-time performance, and collaboration of operational processes. The development of traditional data center operational technologies is closely related to the informatization process, roughly going through three stages: Early stages relied on manual experience and single-point monitoring equipment, risk identification was mostly reactive, systems were isolated, and data was not interconnected; with the popularization of intelligent operational concepts, some subsystems began to be integrated into a unified platform, resulting in linked control based on static rules, but contingency plans were slowly updated and difficult to cope with dynamically changing risks; in recent years, the application of big data and IoT technologies has propelled operational processes into a data-driven stage, enabling real-time risk assessment in localized areas, but core technological bottlenecks have not yet been overcome.
[0003] Current mainstream operation and maintenance (O&M) systems suffer from significant fragmentation and data silos. Systems for environmental monitoring, equipment management, personnel qualifications, and task scheduling are often independently built, with inconsistent data standards and a lack of effective sharing mechanisms and a unified data exchange platform. Data such as O&M personnel qualification files, operational behaviors, and training records are scattered across human resources systems; equipment operating parameters, fault histories, and maintenance records are stored on equipment management platforms; and environmental data such as data center temperature and humidity, and power supply stability are collected by environmental monitoring systems. These data silos between systems result in incomplete O&M risk profiles, hindering a comprehensive understanding of the "human-machine-environment-task" relationship. Furthermore, the difficulty in interconnecting management data from different transportation modes translates into data fragmentation across different O&M stages within the data center. For example, there is a lack of correlation analysis between personnel dynamic behavior data and equipment operating status data, making it impossible to accurately assess the impact of human actions on equipment safety.
[0004] Existing O&M technologies have significant shortcomings in risk perception and control capabilities. Most systems rely on fixed thresholds for risk alerts, failing to dynamically adjust judgment criteria based on real-time environmental changes, equipment status fluctuations, and differences in personnel behavior, resulting in persistently high false alarm and false negative rates. Risk identification is often based on single-dimensional data or simple rules, lacking deep integration and machine learning modeling of multi-dimensional data across "human-machine-environment-task," making it difficult to detect complex, hidden, and transmissible risks, such as the risk of server cluster load imbalance caused by an anomaly in a regional cooling system. Control measures are often discrete and conflicting; each subsystem formulates strategies based on its own management objectives, potentially leading to local optimization at the expense of global deterioration. Furthermore, the lack of real-time quantitative effect evaluation after control commands are issued prevents the formation of a closed loop of "perception-decision-execution-evaluation-optimization." In addition, existing linkages are mostly concentrated on specific equipment or local areas, lacking full-process tracking and collaborative control for complex O&M tasks across data centers and regions. Control commands are mostly directed at management personnel, failing to effectively reach frontline O&M staff, resulting in weak human-machine-environment-cloud collaborative control capabilities.
[0005] Limitations in technical architecture and computing power further restrict the improvement of operation and maintenance (O&M) capabilities. Traditional centralized architectures struggle to handle the massive, high-concurrency, real-time data generated by high-performance data centers, leading to decision-making delays and an inability to meet real-time control requirements. Risk assessment and capability evaluation models are fixed within the system, making it difficult to utilize real-time data streams for rapid iterative optimization. As data center equipment is updated and O&M processes are adjusted, the adaptability of these models gradually declines, failing to accurately match dynamically changing O&M needs. These problems collectively result in traditional O&M systems exhibiting characteristics of being "divide and conquer, lagging, static, and localized," making it difficult to cope with the increasing complexity and dynamism of modern high-performance data centers and the resulting risks and challenges. An innovative O&M system with end-to-end, dynamically interconnected, and interlocking capabilities is urgently needed. Summary of the Invention
[0006] The present invention proposes an automated execution system for the operation and maintenance of high-computing-power data centers to solve the problems mentioned in the prior art.
[0007] To achieve the above objectives, the present invention adopts the following technical solution: an automated operation and maintenance system for high-computing-power data centers, comprising: Data acquisition module: Collects information on the qualifications and behavior of maintenance personnel, operation and fault data of maintenance equipment, and information on the data center environment and maintenance tasks. It collects data in real time through satellite positioning and image recognition technology, and integrates historical and static data from the existing platform system. Operations and maintenance personnel capability assessment module: Combining the collected operations and maintenance personnel data, the module uses weighted analysis of indicators to form the capability status level Cs of operations and maintenance personnel. The indicators include qualification compliance, operation proficiency, fault handling efficiency, safety awareness, training compliance rate, and dynamic behavior standardization. The module establishes a grading standard to divide the capability status into five levels: excellent, good, qualified, need improvement, and unqualified. The grading process integrates real-time behavior data and historical performance data. Operation and maintenance task risk assessment module: Based on operation and maintenance task information, environment information, equipment and status data, calculate the task risk level Rt from four dimensions: task complexity, working environment risk, equipment health risk and security control requirements. The risk level is divided into five levels, and the level division is calibrated with reference to data center security operation and maintenance standards and combined with historical risk records of similar tasks. Task risk level quantification calculation module: Risk level is quantified by weighted summation of risk factors. The calculation expression is Rt=α×Tc+β×Er+γ×Eh+δ×Sr, where Rt is the quantified value of task risk level, α is the weight coefficient of task complexity, Tc is the task complexity score, β is the weight coefficient of work environment risk, Er is the work environment risk score, γ is the weight coefficient of equipment health risk, Eh is the equipment health risk score, δ is the weight coefficient of safety control requirements, and Sr is the safety control requirements score. The quantification calculation result is mapped to five risk levels according to preset thresholds. Two-dimensional interlocking decision module: Constructs a two-dimensional interlocking matrix and rule base of Cs and Rt. The horizontal axis of the matrix represents the capability status level of operation and maintenance personnel (Cs), and the vertical axis represents the risk level of the task (Rt). Different combinations map to corresponding control strategies, including access permissions for operation and maintenance personnel, equipment matching specifications, operation speed limit requirements, configuration of accompanying safety personnel, encryption of inspection frequency, pre-positioning of emergency resources, simplification or enhancement of work processes, etc. The rule base has built-in optimal control solutions for different scenarios and supports dynamic updates based on actual operation and maintenance needs. Real-time linkage control module: During task execution, the system updates the capability status level Cs of operation and maintenance personnel and the task risk level Rt in real time through vehicle terminals, on-site monitoring equipment and environmental sensors. When any dimension exceeds the preset safety range, the system immediately links and adjusts the operation and maintenance resource allocation plan and on-site safety control measures. In high-risk scenarios, the operation and maintenance task is suspended or downgraded, and linkage control instructions are sent to operation and maintenance management personnel in a synchronous manner. Closed-loop update module: Writes back abnormal performance, violations, fault handling results, and capability fluctuation data during the operation and maintenance process to the operation and maintenance personnel's capability profile, automatically triggers targeted safety training and reassessment processes, and matches training content with capability gaps. The reassessment results update the operation and maintenance personnel's capability status level Cs. At the same time, the task risk management effect is fed back to the task risk assessment module to optimize the risk level calculation model and two-dimensional interlocking matrix rules.
[0008] Furthermore, it also includes: The module for calculating the capability status level of operations and maintenance personnel: calculates the core value of capability status level through weighted fusion of indicators. The calculation expression is as follows: Where Cs is the core value of the operation and maintenance personnel's capability status level, i is the evaluation indicator number, and w i For the weight coefficient of the i-th indicator to satisfy w1+w2+w3+w4+w5+w6=1, S i Let k be the base score for the i-th indicator. i D is the dynamic correction coefficient for the i-th indicator. i Let be the real-time fluctuation coefficient of the i-th indicator.
[0009] Data Coding and Cleaning Module: All collected data is standardized and coded using a unified coding rule. The coding includes data type, source identifier, timestamp, and check code. The coded data is stored in four categories: personnel, equipment, environment, and task. At the same time, multi-level filters and big data analysis methods are used to screen and classify the collected data, removing invalid data, false alarm data, and duplicate data.
[0010] The decision optimization module utilizes deep learning algorithms to analyze historical operation and maintenance data, the effectiveness of control strategies, and risk evolution data. It optimizes the rule configuration and indicator weights of the two-dimensional interlocking matrix, predicts the effectiveness of different control strategies under specific Cs and Rt combinations through the trained model, automatically recommends the optimal control scheme, and adaptively adjusts the evaluation indicators and risk factors according to changes in the scale of data center operation and maintenance, equipment type, and operational characteristics.
[0011] Furthermore, it also includes: Safety training linkage module: Based on the shortcomings and abnormal performance of operation and maintenance personnel, training programs are generated. The training content includes operating procedures, safety procedures, fault handling skills, and emergency response procedures. The training formats include online courses, offline hands-on practice, simulation exercises, and case analysis. After the training is completed, a re-evaluation is conducted through theoretical examinations, practical assessments, and emergency drill evaluations. Those who pass the re-evaluation will have their capability status level updated.
[0012] The real-time monitoring module for on-site environment and equipment status is equipped with temperature and humidity sensors, smoke sensors, infrared sensors, vibration sensors, and current and voltage sensors to collect data on the computer room environment and equipment operation in real time. The sensor sampling frequency is dynamically adjusted according to the importance of the equipment and the stability of the environment. The monitoring data is uploaded to the system platform in real time via wireless transmission technology. When the monitoring data exceeds the safety threshold, an early warning is immediately triggered and pushed to the two-dimensional interlock decision module and the real-time linkage control module.
[0013] Furthermore, it also includes: Operation and maintenance resource matching optimization module: Based on the control strategy output by the two-dimensional interlock decision module, optimize the matching scheme between operation and maintenance personnel and equipment. High-risk tasks are prioritized to match high-capability operation and maintenance personnel with stable performance equipment. At the same time, according to the work area and task type, nearby emergency resources and accompanying safety personnel are reasonably dispatched to optimize operation and maintenance routes and work processes.
[0014] Early warning and emergency response module: It sets up early warning mechanisms for capability level, risk level, data anomaly, equipment failure, and environmental change. Different early warning types correspond to different early warning methods and response procedures. Early warning information is released simultaneously through audible and visual alarms, SMS notifications, APP push notifications, and platform pop-ups. The emergency response process is based on the early warning level and scenario type, with pre-set standardized handling steps to coordinate relevant operation and maintenance resources and management personnel to respond.
[0015] Furthermore, it also includes: Model Iteration and Update Module: Regularly collects all data during the operation and maintenance process, including data collection quality assessment results, the effectiveness of control strategies, and risk evolution trajectory. It uses big data analysis methods to identify deviations and deficiencies in the model, optimizes the weights of capability assessment indicators, risk level calculation factors, and two-dimensional interlocking matrix rules, and fully deploys the updated model after verifying its accuracy through a small-scale pilot test.
[0016] Compared with existing technologies, the beneficial effects of this invention are: The system integrates multi-source data through a data acquisition module, breaking down data barriers between environmental monitoring, equipment management, personnel management, and task scheduling. It employs unified coding and purification technologies to standardize data formats, eliminating invalid data and supplementing missing data to ensure data integrity and reliability. This deep integration of multi-dimensional data allows operational risk assessment to move beyond single-stage considerations, enabling a comprehensive understanding of "people-machine-environment-tasks." This avoids the narrow risk perspective caused by data fragmentation, laying a solid foundation for accurate assessment and scientific decision-making.
[0017] The collaborative operation of the operations and maintenance personnel capability assessment module and the task risk assessment module enables precise quantification of personnel capabilities and task risks. The personnel capability status level (Cs) integrates real-time behavioral data and historical performance to comprehensively reflect the actual operational capabilities of operations and maintenance personnel. The task risk level (Rt) comprehensively considers the risk level of the operations and maintenance scenario from multiple dimensions, overcoming the limitations of traditional static thresholds. The matrix and rule base constructed by the two-dimensional interlocking decision module accurately match personnel capabilities with task risks, outputting appropriate control strategies to avoid security risks caused by mismatches between capabilities and risks. Simultaneously, it optimizes resource allocation, improves operational efficiency, and completely changes the drawbacks of traditional local optimization leading to global deterioration.
[0018] The real-time linkage and control module ensures the dynamic adaptability of the operation and maintenance process. It continuously updates personnel capabilities and task risk levels through real-time monitoring data. If the risk exceeds safe limits, it immediately adjusts strategies, suspending or downgrading tasks if necessary, achieving rapid risk response and effective risk control. This dynamic interlocking mechanism makes control measures more targeted and timely, significantly reducing the probability of risk spread and improving the safety and stability of the operation and maintenance process.
[0019] The closed-loop update module and the model iteration update module form a virtuous cycle of continuous optimization. Abnormal data generated during operations and maintenance is written back to personnel capability profiles, automatically triggering personalized security training and reassessment, thus driving continuous improvement in the capabilities of operations and maintenance personnel. Simultaneously, the control effectiveness is fed back to the risk assessment model and interlock matrix, enabling dynamic optimization of models and rules. This allows the system to adapt to dynamic changes such as data center equipment updates and process adjustments, maintaining high-efficiency operation and maintenance control capabilities in the long term.
[0020] Overall, the system constructs a complete operation and maintenance system of "perception-decision-execution-evaluation-optimization" through end-to-end data interoperability, multi-dimensional accurate assessment, dynamic interlocking decision-making, real-time linkage control, and continuous closed-loop optimization. This not only improves the security, accuracy, and efficiency of operation and maintenance, but also enhances the adaptability to complex scenarios and dynamic requirements. It provides comprehensive and highly reliable operation and maintenance support for the stable operation of high-computing-power data centers, and helps the core infrastructure of the digital economy operate safely and efficiently. Attached Figure Description
[0021] Figure 1 This is a schematic block diagram of the automated operation and maintenance system for high-computing-power data centers proposed in this invention. Figure 2 This is a comparative diagram showing the evaluation results of the operational personnel's capability status level in the automated execution system for high-computing-power data center operations and maintenance proposed in this invention. Figure 3 This is a schematic diagram illustrating the dynamic changes of Cs and Rt during the task execution process of the automated execution system for high-computing-power data center operation and maintenance proposed in this invention. Figure 4 This diagram illustrates the comparison of response times for different operation and maintenance scenarios in the automated execution system for high-performance data center operations and maintenance proposed in this invention. Detailed Implementation
[0022] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0023] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.
[0024] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.
[0025] Reference Figures 1 to 4 A high-performance data center operation and maintenance automation system, comprising: Data Acquisition Module: Collects information on maintenance personnel qualifications, violation records, accident records, physical examination results, training records, online operation behavior, fatigue status, and concentration data. Simultaneously, it collects operating parameters, performance test reports, maintenance records, and fault history of maintenance equipment; environmental information such as temperature and humidity, cleanliness, power supply stability, cooling system status, fire protection facility status, network bandwidth utilization, and server load rate of the data center; and task information such as maintenance task type, execution time, work area, equipment correlation, and security level. It uses satellite positioning technology, image recognition technology, multiple types of sensors, and wireless transmission technology to achieve real-time data acquisition and integrates existing integrated management platform and maintenance scheduling system to complete the aggregation of historical and static data. Operations and maintenance personnel capability assessment module: Combining the collected operations and maintenance personnel data, the module uses multi-dimensional indicators to weighted analysis to form the operation and maintenance personnel capability status level Cs. The indicators cover qualification compliance, operation proficiency, fault handling efficiency, safety awareness training compliance rate, and dynamic behavior standardization. The module establishes a grading standard to divide the capability status into five levels: excellent, good, qualified, need improvement, and unqualified. The grading process integrates real-time behavior data and historical performance data. Operation and maintenance task risk assessment module: Based on operation and maintenance task information, environmental information, and equipment status data, the task risk level Rt is calculated from four dimensions: task complexity, operating environment risk, equipment health risk, and security control requirements. The risk level is divided into five levels: low risk, medium risk, relatively high risk, high risk, and extremely high risk. The level division is calibrated with reference to data center security operation and maintenance standards and historical risk records of similar tasks. The two-dimensional interlocking decision module constructs a two-dimensional interlocking matrix Cs and Rt and a rule base. The horizontal axis of the matrix represents the capability status level of operation and maintenance personnel (Cs), and the vertical axis represents the task risk level (Rt). Different combinations map to corresponding control strategies, including access permissions for operation and maintenance personnel, equipment matching specifications, operation speed limits, configuration of accompanying safety personnel, encryption of inspection frequency, pre-positioning of emergency resources, simplification or enhancement of work processes, etc. The rule base has built-in optimal control solutions for different scenarios and supports dynamic updates based on actual operation and maintenance needs. Real-time linkage control module: During task execution, the system updates the capability status level Cs of operation and maintenance personnel and the task risk level Rt in real time through vehicle terminals, on-site monitoring equipment and environmental sensors. When any dimension exceeds the preset safety range, the system immediately links and adjusts the operation and maintenance resource allocation plan and on-site safety control measures. In high-risk scenarios, the system can suspend or downgrade the operation and maintenance task and send linkage control instructions to operation and maintenance management personnel and operation and maintenance personnel in a synchronized manner. Closed-loop update module: Writes back abnormal performance, violations, fault handling results, and capability fluctuation data during operation and maintenance to the capability profile of operation and maintenance personnel, automatically triggers targeted safety training and reassessment process, the training content is accurately matched according to capability gaps, the reassessment results update the capability status level Cs of operation and maintenance personnel, and at the same time feeds back the task risk control effect to the task risk assessment module to optimize the risk level calculation model and two-dimensional interlock matrix rules.
[0026] This invention also includes: The module for accurately calculating the capability status level of operations and maintenance personnel: calculates the core value of capability status level by weighting and fusing multiple dimensions of indicators. The calculation expression is as follows: Cs represents the core numerical value of the operation and maintenance personnel's capability status level, ranging from 0 to 100; i represents the evaluation indicator number, with values from 1 to 6 corresponding to qualification compliance, operational proficiency, fault handling efficiency, security awareness, training compliance rate, and dynamic behavioral standardization, respectively; w iFor the weight coefficient of the i-th indicator to satisfy w1+w2+w3+w4+w5+w6=1, S i k is the base score for the i-th indicator, ranging from 0 to 100. i is the dynamic correction coefficient for the i-th indicator, with a value ranging from -0.2 to 0.3; D i The real-time fluctuation coefficient of the i-th indicator ranges from -1 to 1. The weighted fusion and dynamic correction of multi-dimensional indicators make the capability level more consistent with the real-time operating status of maintenance personnel, providing data support for the accurate matching of management and control strategies.
[0027] Data Coding and Cleaning Module: All collected data is standardized and coded using unified coding rules. The coding includes data type, source identifier timestamp, and check code. The coded data is stored in four categories: personnel, equipment, environment, and task. At the same time, multi-level filters and big data analysis methods are used to screen and classify the collected data, removing invalid data, false alarm data, and duplicate data. Missing data is supplemented using linear interpolation. Through data cleansing, data quality and computational efficiency are improved, providing a reliable data foundation for subsequent capability assessment, risk calculation, and decision-making.
[0028] This invention also includes: Task Risk Level Quantification Module: Risk level quantification is achieved through weighted summation of multi-dimensional risk factors. The calculation expression is Rt=α×Tc+β×Er+γ×Eh+δ×Sr, where Rt is the quantified task risk level value (range 0-100), α is the task complexity weight coefficient (range 0.25-0.35), Tc is the task complexity score (range 0-100), β is the operating environment risk weight coefficient (range 0.25-0.35), Er is the operating environment risk score (range 0-100), γ is the equipment health risk weight coefficient (range 0.2-0.3), Eh is the equipment health risk score (range 0-100), δ is the safety control requirement weight coefficient (range 0.1-0.2), and Sr is the safety control requirement score (range 0-100). The quantification results are mapped to five risk levels according to preset thresholds, making risk assessment more objective and comparable. Weight coefficients: α, β, γ, and δ are the weight coefficients of each risk factor, satisfying α+β+γ+δ=1. This weight allocation is set according to the degree of impact of operation and maintenance risks in high-performance data centers. Task complexity (α) and operating environment risk (β) are the core risk factors in the operation and maintenance process, directly determining the basic risk level of task execution, so a relatively high weight range of 0.25-0.35 is set. Equipment health risk (γ) is the core risk source on the equipment side, playing a key role in the success or failure of operation and maintenance tasks, and a weight range of 0.2-0.3 is set. Security control requirements (δ) are the basic constraints of the operation and maintenance process, serving as a fallback dimension for risk assessment, and a weight range of 0.1-0.2 is set. At the same time, each weight coefficient can be dynamically adjusted according to the operation and maintenance scenario (daily operation and maintenance / emergency operation and maintenance), equipment type (core server / ordinary supporting equipment), and operating area (core computer room / edge auxiliary area) of the high-performance data center to adapt to personalized operation and maintenance needs.
[0029] The scoring criteria, Tc, Er, Eh, and Sr, are all quantified using a 0-100 point scale. The scoring dimensions are aligned with the actual operation and maintenance of high-performance data centers. Task complexity (Tc) is scored from four dimensions: the number of devices involved, the extent of cross-data center / cross-module operations, the complexity of the operation steps, and whether downtime is required. Higher scores indicate more complex task execution processes. Operating environment risk (Er) is scored from five dimensions: data center temperature and humidity stability, power supply continuity, cooling system operation status, fire protection effectiveness, and personnel flow in the work area. Higher scores indicate more uncertainties and risks in the operating environment. Equipment health risk (Eh) is scored from four dimensions: equipment age, historical failure rate, real-time performance parameters (CPU load, memory usage, etc.), and maintenance frequency. Higher scores indicate higher operational risks for the equipment itself. Security control requirements (Sr) are scored from four dimensions: data security level, operation permission level, emergency response standards, and operational process compliance. Higher scores indicate stricter security control standards for task execution.
[0030] Mapping of Quantitative Values to Risk Levels: The Rt quantitative value ranges from 0 to 100 and is the comprehensive risk value after weighting various risk factors. The higher the score, the higher the overall risk level of the operation and maintenance task. The preset mapping thresholds between Rt and the five risk levels are as follows: 0-40 points correspond to low risk, 41-60 points correspond to medium risk, 61-80 points correspond to relatively high risk, 81-90 points correspond to high risk, and 91-100 points correspond to extremely high risk. These thresholds are determined after statistical analysis of the probability of risk occurrence and the degree of loss of similar operation and maintenance tasks in high-computing-power data centers over the past three years. They can also be fine-tuned according to the operation and maintenance scale and business importance of the data center to ensure that the risk level classification matches the actual operation and maintenance risks.
[0031] Parameter value constraints: Each scoring item (Tc, Er, Eh, Sr) has a 0-100 point scale with clearly defined quantitative scoring rules. There are no subjective scoring items, and the values of all parameters are automatically calculated by the system's built-in quantitative scoring model. Combined with objective data from the on-site monitoring equipment of the high-performance data center and the historical operation and maintenance database, the system avoids evaluation bias caused by human intervention and ensures the objectivity, accuracy and repeatability of the quantitative calculation of task risk level.
[0032] The decision-making optimization module utilizes deep learning algorithms to analyze historical operation and maintenance data, the effectiveness of management and control strategies, and risk evolution data. It optimizes the rule configuration and indicator weights of the two-dimensional interlocking matrix, predicts the effectiveness of different management and control strategies under specific Cs and Rt combinations through the trained model, automatically recommends the optimal management and control scheme, and adaptively adjusts the evaluation indicators and risk factors according to changes in data center operation and maintenance scale, equipment type, and operation characteristics to improve the system's adaptability to different scenarios.
[0033] This invention also includes: The safety training linkage module generates personalized training plans based on the shortcomings and abnormal performance of operation and maintenance personnel. The training content covers operating procedures, safety procedures, fault handling skills, emergency response procedures, etc. The training formats include online courses, offline hands-on practice, simulation exercises, and case analysis. After the training is completed, re-evaluation is carried out through theoretical examinations, practical assessments, and emergency drill evaluations. Those who pass the re-evaluation will have their ability status level updated, while those who fail will re-attend targeted reinforcement training to ensure the continuous improvement of the operation and maintenance personnel's capabilities.
[0034] The real-time monitoring module for on-site environment and equipment status deploys multiple types of sensors, including temperature and humidity sensors, smoke sensors, infrared sensors, vibration sensors, and current and voltage sensors, to collect real-time environmental parameters and equipment operation data in the computer room. The sensor sampling frequency is dynamically adjusted according to the importance of the equipment and the stability of the environment, with higher sampling frequencies for important equipment and critical areas than for ordinary areas. The monitoring data is uploaded to the system platform in real time via wireless transmission technology. When the monitoring data exceeds the safety threshold, an early warning is immediately triggered and simultaneously pushed to the two-dimensional interlock decision-making module and the real-time linkage control module.
[0035] This invention also includes: Operation and maintenance resource matching optimization module: Based on the control strategy output by the two-dimensional interlock decision module, the matching scheme between operation and maintenance personnel and equipment is optimized. High-risk tasks are prioritized to match high-capability operation and maintenance personnel with stable performance equipment. At the same time, the nearest emergency resources and accompanying safety personnel are reasonably dispatched according to the work area and task type, the operation and maintenance routes and work processes are optimized, unnecessary waiting time and resource waste are reduced, and operation and maintenance efficiency and safety assurance level are improved.
[0036] Multi-dimensional early warning and emergency response module: It sets up multi-dimensional early warning mechanisms such as capability level early warning, risk level early warning, data anomaly early warning, equipment failure early warning, and environmental change early warning. Different early warning types correspond to different early warning methods and response procedures. Early warning information is simultaneously released through multiple methods such as audible and visual alarms, SMS notifications, APP push, and platform pop-ups. The emergency response process is based on the early warning level and scenario type, with pre-set standardized handling steps to coordinate relevant operation and maintenance resources and management personnel for rapid response, minimizing operation and maintenance risks and losses.
[0037] This invention also includes: Model Iteration and Update Module: Regularly collects all data during the operation and maintenance process, including data collection quality assessment results, the effectiveness of control strategy execution, and risk evolution trajectory. Through big data analysis methods, it identifies deviations and deficiencies in the model, optimizes the weights of capability assessment indicators, risk level calculation factors, and two-dimensional interlocking matrix rules. After verifying the updated model through a small-scale pilot test, it is fully deployed to ensure that the system maintains accurate risk identification and efficient linkage control capabilities in the long term, adapting to the dynamic changes in data center operation and maintenance needs.
[0038] Example 1: Application of daily operation and maintenance scenarios in large-scale cloud data centers This embodiment is applied to a large cloud data center, which contains 10 server room modules and deploys more than 20,000 servers. Daily operation and maintenance involves multiple tasks such as equipment inspection, system upgrades, fault diagnosis, and environmental control. It needs to manage more than 80 operation and maintenance personnel at the same time. The core requirement is to improve the security and efficiency of operation and maintenance through full-process linkage control and reduce the risk of business interruption caused by human error and equipment failure.
[0039] The data acquisition module comprehensively covers all dimensions of data related to "people, machines, environment, and tasks." It collects maintenance personnel qualification information, including professional qualification certificate levels, specialized skill certifications, and data on violations and accidents from the past three years. Physical examination results cover indicators such as cardiopulmonary function, vision, hearing, and reaction speed. Training records include annual safety training, equipment operation training, and emergency drill assessment results and attendance. Online operation behavior is collected via image recognition equipment in the server room, including key press frequency, operation accuracy, and number of violations. Fatigue status and concentration data are monitored in real time via wearable devices. Maintenance equipment operating parameters include server CPU load, memory usage, disk read / write speed, and network port traffic. Performance test reports are quarterly third-party test results. Maintenance records detail the time and content of cleaning, parts replacement, and system optimization. Fault history covers the types of equipment faults, troubleshooting time, and solutions from the past five years. The data center environment information is collected through deployed temperature, humidity, smoke, infrared, vibration, current, and voltage sensors. Temperature and humidity are sampled every minute. Power supply stability is monitored by tracking voltage fluctuations and power outages. Cooling system status includes air conditioning operating power, cooling water flow, and cooling effect. Fire protection facilities are monitored in real-time, with pressure values and alarm functions. Network bandwidth utilization and server load are synchronously obtained through the network management platform. Maintenance task information includes task types categorized as inspection, upgrade, troubleshooting, and control; execution time periods differentiate between weekday peak hours, off-peak hours, and nighttime; work areas clearly indicate data center modules and rack numbers; equipment correlation indicates the number of cross-rack and cross-module devices involved in the task; and security levels are classified into three levels: ordinary, important, and critical. Data acquisition utilizes satellite positioning technology to track the location of on-site maintenance personnel, image recognition technology to capture operational behavior and environmental conditions, multiple types of sensors for accurate physical quantity collection, and wireless transmission technology to ensure real-time data upload. Simultaneously, it integrates the existing integrated management platform and maintenance scheduling system of the data center, compiling historical maintenance data and static equipment archives from the past three years.
[0040] The data encoding and purification module standardizes the collected data. A 16-bit unified encoding rule is adopted: the first four bits identify the data type (0001 for personnel data, 0002 for equipment data, 0003 for environmental data, and 0004 for task data); the middle four bits identify the source, corresponding to different collection devices and systems; and the last eight bits contain a timestamp and checksum. The encoded data is categorized into four main types: personnel, equipment, environment, and task, and stored in a unified data space. Data purification employs a three-tiered filter: the first tier filters obviously abnormal data, such as temperature and humidity values exceeding physical limits; the second tier removes duplicate data, such as identical values collected three times consecutively by the same sensor; and the third tier filters false alarms, such as equipment malfunction alarms caused by temporary sensor failures. Missing data is supplemented using linear interpolation. For example, if maintenance personnel's attention data is missing for a certain period, a supplementary value is calculated based on the data trend over the preceding and following 10 minutes to ensure data integrity. The maintenance personnel capability assessment module generates a capability status level Cs through multi-dimensional indicator weighted analysis. A formula is used... Where i = 1 corresponds to qualification compliance, with a weight w1 of 0.2; i = 2 corresponds to operational proficiency, with a weight w2 of 0.25; i = 3 corresponds to fault handling efficiency, with a weight w3 of 0.2; i = 4 corresponds to safety awareness, with a weight w4 of 0.15; i = 5 corresponds to training compliance rate, with a weight w5 of 0.1; and i = 6 corresponds to dynamic behavioral standardization, with a weight w6 of 0.1. The sum of the weights of all items is 1. Taking a senior operations and maintenance personnel as an example, their basic scores are as follows: Qualification compliance S1: 95; Operation proficiency S2: 92; Fault handling efficiency S3: 90; Safety awareness S4: 88; Training compliance rate S5: 94; Dynamic behavioral standardization S6: 91. The dynamic correction coefficients for qualification compliance (k1) are 0.1, operation proficiency (k2): 0.15, fault handling efficiency (k3): 0.2, safety awareness (k4): 0.08, training compliance rate (k5): 0.05, and dynamic behavioral standardization (k6): 0.12. The real-time fluctuation coefficients for qualification compliance (D1): 0.3, operation proficiency (D2): 0.2, fault handling efficiency (D3): 0.4, safety awareness (D4): 0.1, training compliance rate (D5): 0.05, and dynamic behavioral standardization (D6): 0.25. The calculation yields Cs = 0.2 × 95 × (1 + 0.1 × 0.3) + 0.25 × 92 × (1 + 0.15 × 0.2) + 0.2 × 90 × (1 + 0.2 × 0.4) + 0.15 × 88 × (1 + 0.08 × 0.1) + 0.1 × 94 × (1 + 0.05 × 0.05) + 0.1 × 91 × (1 + 0.12 × 0.25) = 0. 2×95×1.03+0.25×92×1.03+0.2×90×1.08+0.15×88×1.008+0.1×94×1.0025+0.1×91×1.03=19.57+23.69+19.44+13.3056+9.4235+9.373=94.8, corresponding to an excellent ability status level. The grading criteria clearly define excellent as 90-100 points, good as 80-89 points, satisfactory as 70-79 points, needing improvement as 60-69 points, and unsatisfactory as below 60 points. The grading process integrates real-time behavioral data and historical performance data, with historical data accounting for 40% and real-time data accounting for 60%.
[0041] The operation and maintenance task risk assessment module calculates the task risk level Rt based on multi-dimensional data, using the formula Rt=α×Tc+β×Er+γ×Eh+δ×Sr, where α is the task complexity weight (0.3), β is the operating environment risk weight (0.3), γ is the equipment health risk weight (0.25), and δ is the security control requirement weight (0.15). Taking a server cluster system upgrade task as an example, the task complexity Tc is 75 points, involving 500 servers across 3 data center modules, requiring downtime; the operating environment risk Er is 60 points, the current temperature and humidity in the data center are normal, but the upgrade period coincides with a low-traffic period with minimal personnel movement; the equipment health risk Eh is 55 points, the cluster servers have been in operation for 2 years with no major failure history; and the security control requirement Sr is 85 points, the system upgrade involves core business data requiring multiple permission verifications. The calculated Rt = 0.3×75 + 0.3×60 + 0.25×55 + 0.15×85 = 22.5 + 18 + 13.75 + 12.75 = 67, corresponding to a medium risk level. The risk level classification criteria are: low risk (0-40 points), medium risk (41-60 points), relatively high risk (61-80 points), high risk (81-90 points), and extremely high risk (91-100 points). The classification references the "Data Center Security Operation and Maintenance Specifications" and is calibrated based on the risk records of similar tasks in this data center over the past three years.
[0042] The two-dimensional interlocking decision module constructs a two-dimensional interlocking matrix (Cs, Rt) and a rule base. The matrix horizontally represents five capability levels: Excellent, Good, Satisfactory, Needs Improvement, and Unsatisfactory; vertically, it represents five risk levels: Low Risk, Medium Risk, Higher Risk, High Risk, and Extremely High Risk. Different combinations map to clearly defined control strategies. For the combination of Excellent (Cs=94.8) and Medium Risk (Rt=67), the rule base outputs the following control strategies: allowing the maintenance personnel to perform tasks independently; requiring them to carry a portable maintenance terminal and fault diagnosis equipment; limiting the operation speed by requiring server shutdowns to follow a preset procedure with at least 30-second intervals between each step; increasing the frequency of inspections by checking the status of core equipment every 15 minutes during the upgrade process; pre-setting emergency resources by placing backup network switches and power modules near the work area; and strengthening the data backup process by requiring dual backups to be completed and their integrity verified before the upgrade. The rule base supports dynamic updates by administrators based on feedback from maintenance performance. For example, if a certain type of task repeatedly experiences incomplete data backups, mandatory backup verification steps can be added to the corresponding combination strategy.
[0043] The real-time linkage control module continuously monitors data during task execution. By updating Cs and Rt in real time through the maintenance terminal in the computer room, on-site monitoring equipment, and environmental sensors, after the server cluster upgrade task was executed for 1 hour, the wearable device detected a decrease in the maintenance personnel's attention data, the dynamic behavior standardization degree S6 dropped from 91 to 78, and the real-time fluctuation coefficient D6 became -0.3. The Cs value was recalculated as follows: Cs = 0.2×95×(1+0.1×0.3)+0.25×92×(1+0.15×0.2)+0.2×90×(1+0.2×0.4)+0.15×88×(1+0.08×0.1)+0.1×94×(1+0.05×0.05)+0.1×78×(1+0.12×(-0.3))=19.57+23.69+19.44+13.3056+9.4235+7.5732=83.0, and the level dropped to good. The system immediately adjusted its control strategy, sending a rest reminder to the personnel via the maintenance terminal, while dispatching an excellent-level maintenance personnel to assist, increasing the inspection frequency to once every 10 minutes, simultaneously enabling real-time video monitoring of the work area, and pushing status change notifications to maintenance management personnel.
[0044] The closed-loop update module records the fluctuation data in the ability level caused by the decline in focus, writes the abnormal performance of the maintenance personnel back to the ability profile, and automatically triggers a targeted safety training process. The training content focuses on focus management, long-term work skills, and fatigue relief methods, and the training format combines online courses with offline practical exercises. After the training, a re-evaluation is conducted through a theoretical exam and simulated work. The re-evaluation results show that the dynamic behavior standardization degree has recovered to 89, Cs has been recalculated to 89.2, and the level has risen back to good. The updated ability status level is synchronized to the maintenance personnel ability assessment module. At the same time, the effect of this adjustment is fed back to the task risk assessment module, optimizing the real-time monitoring threshold of medium-risk tasks, adjusting the initial control strategy of the good level and medium-risk combination in the two-dimensional interlock matrix, and adding a rule that mandates a 15-minute break for work exceeding 1 hour, forming a closed-loop optimization.
[0045] Table 1 is a comparison table of the daily operation and maintenance effects of large-scale cloud data centers: Table 1
[0046] Table 1 clearly demonstrates the significant advantages of this invention in the daily operation and maintenance scenarios of large-scale cloud data centers. Traditional operation and maintenance systems, due to data silos, cannot fully perceive risks, with a risk identification accuracy rate of only 75%. Control responses rely on manual communication, taking 5 minutes, and the failure rate due to human error is as high as 8.2%. This invention achieves accurate risk identification through multi-dimensional data collection and fusion, increasing the accuracy rate to 96%. The real-time linkage control module ensures response to status changes within 30 seconds, quickly adjusting control strategies. The two-dimensional interlocking matrix achieves precise matching of personnel capabilities with task risks, and combined with closed-loop training, improves personnel capabilities, significantly reducing the human error rate. Simultaneously, the resource matching optimization module improves resource utilization to 92%, significantly enhancing the operation and maintenance efficiency of core equipment, fully meeting the security and efficiency requirements of complex daily operation and maintenance in large-scale cloud data centers, and providing a solid guarantee for business continuity.
[0047] Example 2: Application of emergency operation and maintenance scenarios in edge data centers This embodiment is applied to a city edge data center. The data center mainly serves the local government system and enterprise cloud computing needs. The data center has a small area but dense equipment. Emergency operation and maintenance scenarios include equipment repair after extreme weather such as typhoons and rainstorms, emergency troubleshooting of sudden failures, and emergency expansion of important services. The core requirement is to respond quickly to risks in complex and ever-changing emergency scenarios, and ensure safe and efficient emergency operation and maintenance through cross-linking, so as to reduce secondary accidents and business interruption time.
[0048] The data acquisition module optimizes data collection dimensions and frequency for emergency scenarios. It prioritizes emergency response certifications and extreme environment operation experience for emergency maintenance personnel qualification information; focuses on data related to emergency scenarios from the past year for violation and accident records; adds stress resistance test indicators to physical examination results; extracts emergency drill assessment scores and practical performance from training records; collects online operational behavior in real time using portable image recognition devices; and increases the sampling frequency of fatigue status and concentration data to once every 30 seconds. The sampling frequency of maintenance equipment operating parameters is increased to once every 10 seconds under emergency scenarios, focusing on monitoring equipment damage such as casing integrity, interface looseness, and abnormal operating noises after power-on. Performance test reports provide rapid detection data after a fault occurs; maintenance records extract recent emergency repair content; and fault history prioritizes data on the handling of similar sudden faults. For data center environment information, it adds meteorological data such as wind speed, rainfall, and temperature changes; geological disaster early warning information such as earthquake and flood warnings; power supply stability focuses on monitoring the emergency power supply startup status and backup generator fuel level; and cooling system status focuses on monitoring water ingress and the degree of cooling function recovery after extreme weather. The operation and maintenance task information clearly identifies emergency task types as emergency repair, capacity expansion, and troubleshooting; execution time periods distinguish between normal and extreme weather periods; work areas are marked with hazardous areas and safe passages; equipment correlation clearly identifies the core business equipment involved in the emergency task; and the safety level is uniformly upgraded to critical level. Data collection utilizes satellite positioning technology to track the routes and times of emergency operation and maintenance personnel arriving at the scene; wireless transmission technology prioritizes 5G networks to ensure data transmission in harsh environments; and the emergency command platform and data center operation and maintenance system are integrated to complete the aggregation of emergency resource allocation information and historical emergency response data.
[0049] The data encoding and purification module adapts to the data characteristics of emergency scenarios. The unified encoding rules add emergency data acquisition device type as a source identifier, timestamps are accurate to the second, and checksums employ double encryption to ensure data transmission security. Data purification optimizes filtering rules for high-frequency data collected in emergency scenarios. The first-level filter quickly filters out sensor malfunction data caused by extreme environments, such as abnormally high values from temperature and humidity sensors during heavy rain. The second-level filter removes duplicate alarm data during the emergency process, such as multiple triggers of the same fault. The third-level filter identifies non-critical data under emergency conditions, such as routine operating parameters of fault-free equipment, prioritizing the retention of high-value emergency data. Missing data is supplemented using interpolation methods based on historical emergency scenario data. For example, if sensor damage in a certain area leads to data loss after a typhoon, supplementary values are calculated based on surrounding sensor data and historical environmental change patterns during typhoons.
[0050] The operation and maintenance personnel capability assessment module adjusts the indicator weights for emergency scenarios, using a formula. Where i = 3 corresponds to fault handling efficiency, with its weight w3 adjusted to 0.3; i = 6 corresponds to dynamic behavioral standardization, with its weight w6 adjusted to 0.15; i = 1 corresponds to qualification compliance, with a weight w1 of 0.15; i = 2 corresponds to operational proficiency, with a weight w2 of 0.2; i = 4 corresponds to safety awareness, with a weight w4 of 0.15; and i = 5 corresponds to training pass rate, with a weight w5 of 0.05. The sum of all weights is 1. Taking an emergency maintenance personnel as an example, the basic score S1 for qualification compliance is 85, the basic score S2 for operational proficiency is 88, the basic score S3 for fault handling efficiency is 90, the basic score S4 for safety awareness is 92, the basic score S5 for training pass rate is 80, and the basic score S6 for dynamic behavioral standardization is 86; the dynamic correction coefficient k3 for fault handling efficiency is 0.3, and the dynamic correction coefficients k for the other indicators are... i All are 0.1; the real-time fluctuation coefficient of fault handling efficiency D3 is 0.4, the real-time fluctuation coefficient of dynamic behavior standardization D6 is 0.2, and the real-time fluctuation coefficients of other indicators are D i All are 0.1. The calculated Cs = 0.15 × 85 × (1 + 0.1 × 0.1) + 0.2 × 88 × (1 + 0.1 × 0.1) + 0.3 × 90 × (1 + 0.3 × 0.4) + 0.15 × 92 × (1 + 0.1 × 0.1) + 0.05 × 80 × (1 + 0.1 × 0.1) + 0.15 × 86 × (1 + 0.1 × 0.2) = 12.8625 + 17.776 + 34.2 + 13.858 + 4.04 + 13.272 = 95.998, which is excellent.
[0051] The operation and maintenance task risk assessment module adjusts the risk factor weights for emergency scenarios, using the formula Rt=α×Tc+β×Er+γ×Eh+δ×Sr, where β is the operating environment risk weight, adjusted to 0.35; α is the task complexity weight, with a value of 0.25; γ is the equipment health risk weight, with a value of 0.25; and δ is the safety control requirement weight, with a value of 0.15. Taking the emergency repair task of data center equipment after a typhoon as an example, the task complexity Tc is 80 points, requiring the repair of 10 water-damaged servers involving core government systems; the operating environment risk Er is 90 points, with water accumulation on the data center floor and unstable power supply in some areas; the equipment health risk Eh is 85 points, with the extent of hardware damage to the servers after water damage unknown; and the safety control requirement Sr is 95 points, requiring the protection of government data from loss. The calculated value is Rt = 0.25×80 + 0.35×90 + 0.25×85 + 0.15×95 = 20 + 31.5 + 21.25 + 14.25 = 87, which corresponds to a high risk level.
[0052] The two-dimensional interlocking decision module outputs precise control strategies for high-risk emergency tasks. The combination of an excellent rating (Cs=96) and a high-risk rating (Rt=87) determines the control strategy mapped by the rule base. This strategy requires two-person collaborative operation, with both maintenance personnel being at the excellent rating. Equipment specifications include an emergency repair toolkit, a portable generator, data backup equipment, and waterproof protective gear. Operation speed limits require at least one minute interval between each step of equipment disassembly and power-on testing. One dedicated safety officer is assigned to each team member to monitor the environment and provide risk warnings. Inspection frequency is increased to every 5 minutes for equipment status and environmental data. Emergency resources are pre-positioned near the work site, including an emergency rescue vehicle, backup server, and dehumidification equipment. The workflow emphasizes data backup and waterproof protection, and multi-view video monitoring is enabled throughout the process, with real-time uploads to the emergency command platform.
[0053] The real-time linkage control module dynamically responds to changes in risk during the emergency repair task. Thirty minutes into the operation, environmental sensors detected that the humidity in the computer room rose from 75% to 85%, exceeding the safety threshold. The environmental risk Er was updated to 95 points, and Rt was recalculated as 0.25×80+0.35×95+0.25×85+0.15×95=20+33.25+21.25+14.25=88.75. The risk level remained high, but approached the extremely high risk threshold. The system immediately adjusted its control strategy, sending a humidity exceedance warning to the operators and safety officer, scheduling the on-site dehumidification equipment to operate at full capacity, increasing the inspection frequency to once every 3 minutes, and having the safety officer focus on monitoring humidity changes and the waterproof status of the equipment. Fifteen minutes later, the humidity dropped to 65%, Rt was recalculated to 82, and the system restored its original control strategy, ensuring the safe progress of the operation.
[0054] The closed-loop update module records the risk level fluctuations caused by this sudden humidity change. It writes operational data from personnel in high-humidity environments, such as waterproofing implementation and troubleshooting efficiency, back into their capability profiles. It automatically triggers specialized training on environmental adaptability in emergency scenarios, including extreme environment operation procedures, equipment waterproofing techniques, and emergency risk prediction methods. The reassessment results show that the maintenance personnel's emergency environmental adaptability has improved, with Cs updated to 97.5. Simultaneously, the effectiveness of this emergency control is fed back to the task risk assessment module, optimizing the environmental risk thresholds and weights for high-risk tasks, adjusting the control strategies for excellent and high-risk combinations in the two-dimensional interlock matrix, adding rules for automatic activation of dehumidifiers when humidity exceeds 80%, and improving the emergency maintenance closed loop.
[0055] Table 2 is a comparison table of emergency operation and maintenance effects of edge data centers: Table 2
[0056] Table 2 data highlights the core value of this invention in emergency operation and maintenance scenarios for edge data centers. Traditional emergency operation and maintenance relies on manual experience and static contingency plans, resulting in emergency response times of up to 40 minutes, a fault repair success rate of only 72%, a secondary accident rate as high as 10.5%, and core business interruption duration exceeding 3 hours. This invention achieves rapid risk identification and accurate assessment through an emergency scenario-adaptive data acquisition and purification module; targeted control strategies output by the two-dimensional interlocking decision module ensure reasonable allocation of emergency resources and safe operation processes; a real-time linkage control module dynamically responds to changes in the environment and personnel status, effectively avoiding secondary accidents; and a closed-loop update module accelerates the improvement of operation and maintenance personnel's emergency response capabilities. Ultimately, the emergency response time is shortened to 15 minutes, the fault repair success rate is increased to 95%, and the core business interruption duration is controlled within 45 minutes, fully meeting the rapid response and security requirements of emergency operation and maintenance in edge data centers, and providing strong support for the continuity of local government and enterprise businesses.
[0057] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. An automated operation and maintenance system for high-computing-power data centers, characterized in that: include: Data acquisition module: Real-time collection of information on maintenance personnel, equipment, tasks and environment through satellite positioning and image recognition technology, integration of existing platform systems, and aggregation of historical and static data; Operations and maintenance personnel capability assessment module: Combines the collected operations and maintenance personnel data to form the operation and maintenance personnel capability status level Cs, establishes a grading standard, and divides the capability status into levels; Operation and maintenance task risk assessment module: Based on operation and maintenance task information, environmental information and equipment status data, calculate the task risk level Rt from four dimensions: task complexity, operating environment risk, equipment health risk and safety control requirements; Task risk level quantification calculation module: Risk level is quantified by weighted summation of risk factors. The calculation expression is Rt=α×Tc+β×Er+γ×Eh+δ×Sr, where Rt is the quantified value of task risk level, α is the weight coefficient of task complexity, Tc is the task complexity score, β is the weight coefficient of work environment risk, Er is the work environment risk score, γ is the weight coefficient of equipment health risk, Eh is the equipment health risk score, δ is the weight coefficient of safety control requirements, and Sr is the score of safety control requirements. Two-dimensional interlocking decision module: Constructs a two-dimensional interlocking matrix and rule base of Cs and Rt. The horizontal axis of the matrix represents the capability status level of the operation and maintenance personnel, Cs, and the vertical axis represents the task risk level, Rt. Real-time linkage control module: Cs and Rt are updated in real time during task execution, and immediate linkage adjustment is made when any dimension exceeds the preset safety range; Closed-loop update module: Writes data from the operation and maintenance process back to the operation and maintenance personnel's competency profile, automatically triggers targeted security training and reassessment processes, updates Cs with the reassessment results, and feeds back the task risk control effect to the task risk assessment module.
2. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: The module for calculating the capability status level of operations and maintenance personnel: calculates the core value of capability status level through weighted fusion of indicators. The calculation expression is as follows: Where Cs is the core value of the operation and maintenance personnel's capability status level, i is the evaluation indicator number, and w i For the weight coefficient of the i-th indicator to satisfy w1+w2+w3+w4+w5+w6=1, S i Let k be the base score for the i-th indicator. i D is the dynamic correction coefficient for the i-th indicator. i Let be the real-time fluctuation coefficient of the i-th indicator.
3. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: Data Coding and Cleaning Module: All collected data is standardized and coded using a unified coding rule. The coding includes data type, source identifier, timestamp, and check code. The coded data is stored in four categories: personnel, equipment, environment, and task. At the same time, multi-level filters and big data analysis methods are used to screen and classify the collected data, removing invalid data, false alarm data, and duplicate data.
4. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: The decision optimization module utilizes deep learning algorithms to analyze historical operation and maintenance data, the effectiveness of control strategies, and risk evolution data. It optimizes the rule configuration and indicator weights of the two-dimensional interlocking matrix, predicts the effectiveness of different control strategies under specific Cs and Rt combinations through the trained model, automatically recommends the optimal control scheme, and adaptively adjusts the evaluation indicators and risk factors according to changes in the scale of data center operation and maintenance, equipment type, and operational characteristics.
5. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: Safety training linkage module: Based on the shortcomings and abnormal performance of operation and maintenance personnel, training programs are generated. The training content includes operating procedures, safety procedures, fault handling skills, and emergency response procedures. The training formats include online courses, offline hands-on practice, simulation exercises, and case analysis. After the training is completed, a re-evaluation is conducted through theoretical examinations, practical assessments, and emergency drill evaluations. Those who pass the re-evaluation will have their capability status level updated.
6. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: The real-time monitoring module for on-site environment and equipment status is equipped with temperature and humidity sensors, smoke sensors, infrared sensors, vibration sensors, and current and voltage sensors to collect data on the computer room environment and equipment operation in real time. The sensor sampling frequency is dynamically adjusted according to the importance of the equipment and the stability of the environment. The monitoring data is uploaded to the system platform in real time via wireless transmission technology. When the monitoring data exceeds the safety threshold, an early warning is immediately triggered and pushed to the two-dimensional interlock decision module and the real-time linkage control module.
7. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: Operation and maintenance resource matching optimization module: Based on the control strategy output by the two-dimensional interlock decision module, optimize the matching scheme between operation and maintenance personnel and equipment. High-risk tasks are prioritized to match high-capability operation and maintenance personnel with stable performance equipment. At the same time, according to the work area and task type, nearby emergency resources and accompanying safety personnel are reasonably dispatched to optimize operation and maintenance routes and work processes.
8. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: Early warning and emergency response module: It sets up early warning mechanisms for capability level, risk level, data anomaly, equipment failure, and environmental change. Different early warning types correspond to different early warning methods and response procedures. Early warning information is simultaneously released through sound and light alarms, SMS notifications, and APP push platform pop-ups. The emergency response procedure is based on the early warning level and scenario type, with pre-set standardized handling steps to coordinate relevant operation and maintenance resources and management personnel to respond.
9. The automated operation and maintenance system for high-computing-power data centers according to claim 1, characterized in that, Also includes: Model Iteration and Update Module: Regularly collects all data during the operation and maintenance process, including data collection quality assessment results, the effectiveness of control strategies, and risk evolution trajectory. It uses big data analysis methods to identify deviations and deficiencies in the model, optimizes the weights of capability assessment indicators, risk level calculation factors, and two-dimensional interlocking matrix rules, and fully deploys the updated model after verifying its accuracy through a small-scale pilot test.
Citation Information
Patent Citations
Intelligent operation and maintenance method fusing multi-modal data and active learning
CN120198106A
Intelligent control method and system for special equipment based on Internet of Things
CN121302341A
Server operation and maintenance method and electronic equipment
CN121436967A
Operation and maintenance task processing method, storage medium, and electronic device
WO2025141423A1
Cited By
Artificial intelligence driven cloud service intelligent operation and maintenance security risk dynamic evaluation method and system
CN122268680A