High-computing-power data center environment perception linkage operation and maintenance system

By using the high-performance data center environment perception and linkage operation and maintenance system, combined with biometric and IoT technologies, dynamic assessment and control of operation and maintenance personnel's capabilities and task risks have been achieved. This has solved the problems of decision-making lag and resource scheduling in the operation and maintenance of high-performance data centers, and improved the security and efficiency of operation and maintenance.

CN121903585APending Publication Date: 2026-04-21NINGBO SIHONG ELECTRICAL APPLIANCE IND
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NINGBO SIHONG ELECTRICAL APPLIANCE IND
Filing Date
2026-03-26
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

The operation and maintenance of high-performance computing data centers lacks a fully intelligent management and control system. Operation and maintenance decisions rely on static rules and fail to take into account the real-time status of operation and maintenance personnel and the dynamic changes in tasks. This results in incomplete risk assessment, limited training effectiveness, inefficient resource scheduling, and difficulty in coping with complex and ever-changing operating scenarios.

Method used

The system employs a high-performance data center environment perception and linkage operation and maintenance system. It collects data from operation and maintenance personnel through biometrics and sensor technology, acquires task information by combining IoT sensor networks, uses machine learning algorithms to assess the capabilities of operation and maintenance personnel and task risks, and constructs a two-dimensional interlocking matrix and rule base for dynamic control. This enables full-process data sharing and personalized training, and supports cross-regional collaborative control and data security.

Benefits of technology

It enables multi-dimensional quantification of operation and maintenance personnel capabilities and precise matching of task risks, dynamically adjusts operation and maintenance plans, improves the security, accuracy and efficiency of operation and maintenance, and meets the stable operation requirements of high-computing-power data centers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121903585A_ABST
    Figure CN121903585A_ABST
Patent Text Reader

Abstract

The invention discloses a high-computing-power data center environment perception linkage operation and maintenance system, and relates to the technical field of data center operation and maintenance automatic control and regulation and control. Collecting and transmitting data of qualification, operation, health, training and the like; the operation and maintenance task information acquisition module acquires structured information such as task type equipment state environment parameters; generating an operation and maintenance personnel capability level Cs and a task risk level Rt; the two-dimensional interlocking matrix and the rule base are matched with a management and control strategy; the operation and maintenance scheme is adjusted in real time, and tasks are stopped or degraded if necessary; personalized courses are pushed, and Cs is updated; the classified storage data supports multi-dimensional query; and the display and operation requirements of the management end and the operation and maintenance personnel end are adapted. According to the method, a data barrier is broken, and dynamic and accurate matching between the capability of operation and maintenance personnel and task risks is realized; through full-process closed-loop regulation and control and personalized training, the operation and maintenance safety, accuracy and efficiency are improved, cross-regional collaborative optimization of resource configuration is realized, compliance and safety guarantee are in place, and the operation and maintenance requirements of a high-computing-power data center are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of automatic control and intelligent regulation technology for data center operation and maintenance, and particularly to an environmental perception and linkage operation and maintenance system for high-computing-power data centers. Background Technology

[0002] As a core infrastructure of the digital economy, high-performance computing data centers handle massive data storage and high-performance computing tasks. Their equipment clusters are large-scale, with dense core components and high operating loads, placing extremely high demands on the accuracy, real-time performance, and security of operations and maintenance (O&M) work. Currently, O&M of high-performance computing data centers largely relies on traditional models and has not yet formed a fully intelligent management and control system, resulting in prominent data silos. O&M personnel's qualification information, operation records, health data, and training files are scattered across different management systems. Data such as equipment operating status, operating environment parameters, and task attributes lack a unified integration mechanism. Inconsistent data standards and poor data sharing across systems lead to a lack of comprehensive data support for O&M decisions, making it difficult to accurately characterize the capabilities of O&M personnel and the risk level of tasks.

[0003] Current operational risk control measures suffer from significant lag and bias. They largely rely on static rules and fixed thresholds for risk assessment, failing to dynamically adjust evaluation standards based on factors such as the real-time status of operations personnel, dynamic task changes, and environmental fluctuations. Operations personnel capability assessments depend heavily on periodic static evaluations, neglecting dynamic factors like real-time operational performance and fatigue levels, resulting in insufficient alignment between capability levels and actual capabilities. Task risk assessments focus solely on the equipment's condition, failing to consider multi-dimensional factors such as operational complexity, environmental stability, and historical failures, leading to incomplete risk profiles. Control measures lack a systematic approach, often focusing on localized optimizations that may introduce new operational risks. For example, assigning low-capability personnel to high-risk tasks to meet deadlines, or failing to adjust operational plans according to environmental changes, can lead to equipment failures and data loss.

[0004] The lack of a closed-loop mechanism for operation and maintenance management, coupled with a disconnect between training and actual operational needs, results in limited training effectiveness. Existing training is largely based on standardized courses, failing to develop personalized solutions for abnormal operations and skill gaps among operations personnel. Issues such as violations and improper fault handling during operations are not promptly recorded in personnel competency files, hindering targeted retraining and reassessment. Furthermore, cross-regional resource scheduling is inefficient. When a single data center faces high-risk tasks and lacks suitable local personnel, it cannot quickly coordinate resources from other regions, leading to low operational efficiency. These problems make it difficult for high-performance data center operations to cope with complex and ever-changing operational scenarios, resulting in persistently high operational risks and hindering the stable and efficient operation of data centers. Summary of the Invention

[0005] The present invention proposes a high-computing-power data center environment awareness and linkage operation and maintenance system to solve the problems mentioned in the prior art.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: a high-computing-power data center environment awareness and linkage operation and maintenance system, comprising: Operation and maintenance personnel data collection module: It is equipped with a qualification certification collection unit, an operation behavior monitoring unit, a health status detection unit and a training record integration unit. Through biometrics and sensor technology, it collects qualification, operation, health and training data, and after standardized conversion and verification, it transmits the data to the core processing module twice per second. Operation and maintenance task information collection module: Through the Internet of Things sensor network and the device management system interface, it obtains information on data center operation and maintenance task types, equipment importance, and environmental parameters, and synchronously collects time window constraints to form a structured task information package for transmission to the core processing module. Operations and maintenance personnel capability level assessment module: Deeply processes the data collected by operations and maintenance personnel, removes anomalies through machine learning algorithms, quantifies the score from five dimensions of qualification and compliance, and dynamically adjusts the weights in combination with real-time status to generate an operations and maintenance personnel capability status level Cs from 1 to 5. Task risk level calculation module: The analytic hierarchy process is used to analyze the operation and maintenance task information. It integrates factors such as equipment importance, operation complexity, environmental stability, task urgency, and historical failure frequency, quantifies the weight of each factor and performs weighted calculation to generate the task risk level Rt. The level is divided into 1 to 5 levels, corresponding to risks from low to high. Two-dimensional interlock matrix and rule base module: Constructs a two-dimensional interlock matrix Cs and Rt, and presets control policies corresponding to different combinations of capability levels and risk levels, including access permissions for operation and maintenance personnel, target device matching type, single operation and maintenance duration limit, number of auxiliary operation and maintenance personnel, and operation and maintenance tool adaptation standards. The rule base supports dynamic updating of mapping relationships based on actual operation and maintenance results. Dynamic control execution module: Receives real-time data on the capability level and task risk level of operation and maintenance personnel, matches and executes control strategies according to a two-dimensional interlock matrix, continuously monitors changes in Cs and Rt during task execution, and immediately adjusts the operation and maintenance plan when any dimension exceeds the preset threshold. Targeted training trigger module: The system automatically writes back the operation and maintenance personnel's competency profile with abnormalities such as violations and improper fault handling, matches HSE training courses to generate personalized plans and pushes them, and triggers competency reassessment after training is completed to update the Cs level; Human-Computer Interaction Module: The interface between the system configuration management workstation and personnel mobile terminals displays Cs level, Rt level, and control policy information in real time. It supports administrators in adjusting rules and setting thresholds, and maintenance personnel in viewing files, receiving notifications, and receiving feedback.

[0007] Furthermore, it also includes: The module for accurately calculating the capability status of operations and maintenance personnel is as follows: This module optimizes the calculation logic of the capability level assessment module for operations and maintenance personnel. It quantifies capability status through a multi-dimensional weighted summation algorithm and generates the capability status score of operations and maintenance personnel using the formula Cs=w1A+w2B+w3C+w4D+w5E. Here, Cs is the capability status score, w1 is the qualification compliance weight coefficient, A is the qualification compliance standardization score, w2 is the operation standardization weight coefficient, B is the operation standardization standardization score, w3 is the fault handling capability weight coefficient, C is the fault handling capability standardization score, w4 is the health adaptability weight coefficient, D is the health adaptability standardization score, w5 is the training compliance rate weight coefficient, and E is the training compliance rate standardization score. Each weight coefficient is trained and optimized using historical operation and maintenance data from a high-performance data center and satisfies w1+w2+w3+w4+w5=1. The higher the score, the higher the capability level, thus achieving a quantitative assessment of the capability status of operations and maintenance personnel.

[0008] The dynamic risk warning module for operation and maintenance tasks dynamically adjusts the risk level by monitoring environmental changes, equipment status fluctuations, and operation and maintenance progress in real time during task execution. When the ambient temperature and humidity exceed the safe range, equipment alarms occur suddenly, or operation and maintenance delays exceed the preset threshold, the risk level upgrade process is automatically initiated, and warning information is pushed to the management terminal and operation and maintenance personnel terminals simultaneously. The control strategy is adjusted in combination with the current CS level, and measures such as adding linkage of environmental control equipment and dispatching more technical support personnel to shorten the operation and maintenance interval are implemented.

[0009] Furthermore, it also includes: Operation and maintenance management strategy matching module: Based on a two-dimensional interlocking matrix and rule base, it optimizes the accuracy of strategy mapping through a fuzzy matching algorithm. When Cs and Rt are at the level threshold, the strategy matching score is calculated by formula, where M is the strategy matching score; α is the capability level influence coefficient, β is the risk level influence coefficient, and γ is the normalization coefficient. It also provides an expected effect analysis report for each strategy.

[0010] The fatigue monitoring and intervention module for maintenance personnel analyzes fatigue levels by measuring continuous maintenance duration, operation response speed, frequency of erroneous operations, and blink frequency (biological characteristics). It sets three-level fatigue warning thresholds and records fatigue data to an individual's competency profile, serving as an important basis for health suitability scoring.

[0011] Furthermore, it also includes: High-performance computing equipment operation and maintenance adaptation and optimization module: Establish a database of equipment operation and maintenance characteristics, store operation and maintenance specifications, risk point distribution, and personnel capability requirements for different equipment. When the task risk level Rt=5, automatically match the equipment-specific operation and maintenance tool list, operation step guide, and emergency response plan from the database, configure equipment operation permissions in conjunction with the operation and maintenance personnel's Cs level, and simultaneously link with the equipment management system to lock non-operation and maintenance areas.

[0012] HSE training personalized customization module: Based on abnormal data, a course-abnormal type association mapping library is built. Corresponding courses are matched according to violations and improper handling of faults. Video teaching mode is adopted, training assessment data is recorded, and the pass rate is updated and a capability reassessment is triggered after passing the test, forming a closed loop.

[0013] Furthermore, it also includes: Cross-regional operation and maintenance collaborative control module: Real-time sharing of data, task information and control strategies of operation and maintenance personnel in different regions is realized through cloud communication technology. When a high-risk task Rt=5 occurs in a single data center and there are no matching high-capability level Cs≥4 operation and maintenance personnel locally, the module automatically dispatches suitable personnel from other regions to generate a cross-regional operation and maintenance plan, and synchronously updates the records of data storage modules in each region to realize the optimized configuration and collaborative control of operation and maintenance resources in each data center.

[0014] Data security and compliance audit module: End-to-end encryption is used to protect the transmission and storage of sensitive information. Blockchain is used to prevent tampering and ensure traceability of operation logs. Hierarchical access permissions are set, and compliance audit reports containing core indicators of qualification compliance rate are generated regularly.

[0015] System self-learning optimization module: Collects historical operation and maintenance data, strategy execution effects and training feedback, builds labeled datasets, realizes reinforcement learning algorithm iterative optimization evaluation, calculation and matching model, dynamically adjusts weights and rules, and improves system adaptability, recognition and control capabilities.

[0016] Compared with existing technologies, the beneficial effects of this invention are: The system breaks down data barriers between different modules and systems through a unified data collection and integration mechanism. It comprehensively gathers data on the qualifications, operations, health, and training of maintenance personnel, as well as task information, environmental parameters, and equipment status data. This enables standardized data processing and efficient sharing, providing a solid data foundation for accurate assessment and control, and solving the pain points of fragmented traditional maintenance data and insufficient decision-making basis.

[0017] The dynamic assessment capability has been significantly improved. The operator capability level assessment module combines static qualifications and real-time status to quantify the capability level (Cs) from multiple dimensions. The task risk level calculation module generates the risk level (Rt) through comprehensive multi-factor analysis. A two-dimensional interlocking matrix and rule base achieve precise matching between the two, making management strategies more targeted. The system abandons the static threshold judgment mode and updates Cs and Rt in real time during task execution. Once the threshold is exceeded, the operation and maintenance plan is immediately adjusted, and the task is suspended or downgraded when necessary. This achieves dynamic risk prevention and precise control, avoiding systemic risks caused by local optimization.

[0018] Personalized training and a closed-loop optimization mechanism continuously improve operational capabilities. The system automatically writes back anomalies to personnel competency profiles, accurately matching HSE training courses to form a closed loop of "anomaly detection - customized training - competency reassessment," significantly enhancing the relevance and effectiveness of training and helping operational personnel address skill gaps. The cross-regional collaborative control module enables unified scheduling of resources across multiple data centers, optimizing resource allocation efficiency, while the data security and compliance audit module ensures the security of sensitive information and meets regulatory compliance requirements.

[0019] The system's self-learning and optimization capabilities continuously enhance its operational adaptability. By continuously collecting historical data and iteratively optimizing evaluation models and control rules, it gradually improves its adaptability to complex operational scenarios and the accuracy of risk identification. Overall, the system achieves interlocking and coordinated operation across the entire process of operational personnel assessment, dynamic monitoring, HSE training, on-site risk control, and resource scheduling. This drives the transformation of high-performance data center operations from a static, localized, and manually-driven model to a dynamic, global, and intelligently-driven model, significantly improving operational security, accuracy, and efficiency, and providing strong technical support for the stable operation of data centers. Attached Figure Description

[0020] Figure 1 This is a schematic block diagram of the high-computing-power data center environment perception and linkage operation and maintenance system proposed in this invention; Figure 2 This is a schematic diagram illustrating the changes in personnel capability levels during the operation and maintenance cycle of the high-computing-power data center environment perception and linkage operation and maintenance system proposed in this invention. Figure 3 This is a schematic diagram illustrating the accuracy of different evaluation modules in the high-computing-power data center environment perception and linkage operation and maintenance system proposed in this invention. Figure 4 This is a schematic diagram illustrating the contribution of operation and maintenance data features to the high-computing-power data center environment perception and linkage operation and maintenance system proposed in this invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," and "counterclockwise," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0023] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified. Furthermore, the terms "installed," "connected," and "linked" should be interpreted broadly; for example, they may refer to a fixed connection, a detachable connection, or an integral connection; they may refer to a mechanical connection or an electrical connection; they may refer to a direct connection or an indirect connection through an intermediate medium; and they may refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances. The invention will now be described in further detail with reference to the accompanying drawings.

[0024] Reference Figures 1 to 4 A high-performance data center environment awareness and collaborative operation and maintenance system, comprising: Operations and maintenance personnel data collection module: Equipped with a qualification certification collection unit, an operation behavior monitoring unit, a health status detection unit, and a training record integration unit, it synchronously collects the operation and maintenance personnel's qualification certificate information, violation operation records, operation and maintenance failure events, annual physical examination data, HSE, training qualification results, real-time operation logs, continuous operation and maintenance duration, and mental status data through biometric technology, operation log capture tools, fatigue sensors, and compliance detection terminals. All data is transmitted to the core processing module at a frequency of 2 times per second after standardization conversion and preliminary verification. Operation and maintenance task information collection module: It acquires information on the specific type of operation and maintenance tasks of large computing power data centers, the importance level of target equipment, the complexity of operation and maintenance operations, the temperature and humidity of the working environment, the stability of power supply, the service life of equipment, the work site, HSE, and risk points through IoT sensor networks, equipment management system interfaces, and environmental monitoring terminals. It also collects the time window for task execution, fault tolerance threshold, and related equipment dependency constraints, and forms a structured task information package that is transmitted to the core processing module. Operations and maintenance personnel capability level assessment module: Deeply process the collected operations and maintenance personnel data, remove abnormal data through machine learning algorithms, and quantify the scores according to five dimensions: qualification compliance, operation standardization, fault handling ability, health adaptability, and training compliance rate. Combine the real-time status to dynamically adjust the weights and generate the operations and maintenance personnel capability status level Cs, which is divided into 1 to 5 levels, corresponding to capabilities from low to high. Task risk level calculation module: The hierarchical analysis method is used to perform multi-factor analysis on operation and maintenance task information. It comprehensively considers factors such as equipment importance, operation complexity, environmental stability, task urgency, and historical failure frequency, quantifies the weight of each factor, and performs weighted calculation to generate task risk level Rt. The level is divided into 1 to 5 levels, corresponding to risks from low to high. Two-dimensional interlock matrix and rule base module: Constructs a (Cs,Rt) two-dimensional interlock matrix, presets control policies corresponding to different combinations of capability levels and risk levels, including access permissions for operation and maintenance personnel, target device matching type, single operation and maintenance duration limit, number of auxiliary operation and maintenance personnel, and operation and maintenance tool adaptation standards. The rule base supports dynamic updating of mapping relationships based on actual operation and maintenance results. Dynamic control and execution module: Receives real-time data on the capability level and task risk level of operation and maintenance personnel, matches and executes control strategies according to a two-dimensional interlock matrix, continuously monitors changes in Cs and Rt during task execution, and immediately adjusts the operation and maintenance plan when any dimension exceeds the preset threshold, including extending the operation and maintenance interval, dispatching more technical backbones, changing to low-risk operation periods, activating backup equipment, and suspending or downgrading the current operation and maintenance task when necessary. Targeted training trigger module: Automatically writes back to the operation and maintenance personnel's ability file any abnormal behavior such as violation of operation and maintenance procedures, improper handling of faults, or failure to meet the required ability level. Matches the corresponding HSE training courses according to the type of abnormality, generates a personalized training plan and pushes it to the operation and maintenance personnel's terminal. After the training is completed, the capability re-evaluation process is automatically triggered to update the Cs level. Data storage and management module: It adopts a dual mode of distributed storage architecture and local encrypted caching to classify and store data of operation and maintenance personnel, operation and maintenance task information, capability level assessment results, risk level calculation data, control strategy execution records, training files, and operation logs. It supports multi-dimensional query, export, and traceability by personnel, ID task number, time range, capability level, and risk level. The data storage cycle covers data center operation and maintenance audit and compliance requirements. Human-Computer Interaction Module: Configures the interaction interface between the operation and maintenance management workstation and the mobile terminal of operation and maintenance personnel, displays the operation and maintenance personnel's Cs level, task Rt level, control policy details, operation and maintenance progress, and abnormal warning information in real time. It supports managers to adjust interlock matrix rules, set warning thresholds, and enter evaluation opinions. It also supports operation and maintenance personnel to view their personal capability profiles, receive training notifications, and provide feedback on operation and maintenance status.

[0025] This invention also includes: The module for accurately calculating the capability status level of operations and maintenance personnel: This module optimizes the calculation logic of the capability level assessment module for operations and maintenance personnel. It quantifies capability status through a multi-dimensional weighted summation algorithm and generates the capability status level score of operations and maintenance personnel using the formula Cs=w1A+w2B+w3C+w4D+w5E. Here, Cs is the capability status level score, ranging from 0-100 points corresponding to levels 1-5; w1 is the qualification compliance weight coefficient; A is the qualification compliance standardization score, ranging from 0-100 points; w2 is the operational standardization weight coefficient; and B is the operational standardization score. The score ranges from 0 to 100. w3 is the weighting coefficient for fault handling capability, C is the standardized score for fault handling capability (0-100), w4 is the weighting coefficient for health adaptability, D is the standardized score for health adaptability (0-100), w5 is the weighting coefficient for training compliance rate, and E is the standardized score for training compliance rate (0-100). Each weighting coefficient is optimized through training with historical data from the operation and maintenance of high-performance data centers and satisfies w1+w2+w3+w4+w5=1. The higher the score, the higher the capability level, thus achieving a precise quantitative assessment of the capability status of operation and maintenance personnel.

[0026] Dynamic Risk Warning Module for Operation and Maintenance Tasks: This module enhances the real-time and comprehensiveness of the task risk level calculation module. By monitoring environmental changes, equipment status fluctuations, and operation and maintenance progress during task execution in real time, it dynamically corrects the risk level. When it detects that the ambient temperature and humidity exceed the safe range, equipment alarms suddenly occur, or operation and maintenance delays exceed preset thresholds, it automatically initiates the risk level escalation process and simultaneously pushes warning information to the management terminal and operation and maintenance personnel terminals. Based on the current CS level, it adjusts the control strategy, increases the linkage of environmental control equipment, dispatches more technical support personnel, and shortens the operation and maintenance interval to avoid potential risks in the operation and maintenance process in advance.

[0027] This invention also includes: The intelligent matching module for operation and maintenance management strategies: This module is based on a two-dimensional interlocking matrix and a rule base. It optimizes the accuracy of strategy mapping through a fuzzy matching algorithm. When Cs and Rt are at the level threshold, it uses a formula... The system calculates the strategy matching score, where M is the strategy matching score, ranging from 0 to 100; α is the capability level influence coefficient, β is the risk level influence coefficient, and γ is the normalization coefficient. The higher the score, the stronger the strategy adaptability. The system automatically selects the top 3 management strategies for managers to choose from, and provides an expected effect analysis report for each strategy, including the risk reduction, operational efficiency, impact on cost consumption, and improving the scientific nature of management decisions.

[0028] Intelligent monitoring and intervention module for maintenance personnel fatigue: This module integrates fatigue monitoring data from the maintenance personnel data collection module. It analyzes fatigue levels from multiple dimensions, including continuous maintenance duration, operation response speed, frequency of erroneous operations, blink frequency (biological characteristics), and heart rate. It sets three fatigue warning thresholds. When fatigue is mild, it pushes rest reminders and eye relaxation guidance. When fatigue is moderate, it automatically adjusts the priority of maintenance tasks and extends rest intervals. When fatigue is severe, it immediately triggers a maintenance task pause command. At the same time, it records fatigue status data to personal ability profiles, which serves as an important basis for health suitability scoring and reduces the risk of maintenance failures caused by fatigue.

[0029] This invention also includes: High-performance computing equipment operation and maintenance adaptation optimization module: This module establishes an equipment operation and maintenance characteristic database for the operation and maintenance needs of core equipment such as cooling system power supply equipment of high-performance computing data center server clusters. It stores the operation and maintenance operation specifications, risk point distribution, and personnel capability requirements for different equipment. When the task risk level Rt=5, it automatically matches the equipment-specific operation and maintenance tool list, operation step guide, and emergency response plan from the database. Combined with the Cs level configuration of operation and maintenance personnel, it restricts low-level personnel from accessing critical operations of core equipment. At the same time, it links with the equipment management system to lock non-operation and maintenance areas to avoid equipment failures caused by misoperation.

[0030] HSE Training Personalized Customization Module: This module builds a mapping library between training courses and anomaly types based on the abnormal performance data of the targeted training trigger module. For violations, it matches safety standard practice courses; for improper fault handling, it matches technical problem-solving case courses; and for low ability levels, it matches basic skills improvement courses. It adopts a training mode that combines video teaching, virtual simulation operation, and online assessment. It records training duration, learning progress, and assessment results. After passing the assessment, it automatically updates the training compliance rate score and triggers the ability reassessment process, forming a closed loop of anomaly detection, customized training, and ability improvement.

[0031] Cross-regional operation and maintenance collaborative control module: This module supports the overall scheduling of operation and maintenance resources of multiple data centers. Through cloud communication technology, it realizes real-time sharing of data, task information and control policies of operation and maintenance personnel in different regions. When a high-risk task Rt=5 occurs in a single data center and there are no matching high-capability level Cs≥4 operation and maintenance personnel locally, it automatically dispatches suitable personnel from other regions to generate a cross-regional operation and maintenance plan, including itinerary arrangement, task handover process, remote operation permission configuration, and synchronously updates the relevant records of data storage modules in each region, so as to realize the optimized configuration and collaborative control of operation and maintenance resources of multiple data centers.

[0032] This invention also includes: Data Security and Compliance Audit Module: This module uses end-to-end encryption algorithms to encrypt the transmission and storage of sensitive information such as data collection evaluation results and control strategies. It uses blockchain technology to ensure the tamper-proof and traceability of operation logs, sets up a hierarchical data access permission mechanism, and allows administrators to view only data within their authorized scope. It also automatically generates compliance audit reports on a regular basis, including core indicators such as the compliance rate of operation and maintenance personnel qualifications, training compliance rate, risk control execution rate, and timeliness of anomaly handling, thus meeting the compliance and regulatory requirements for the operation and maintenance of high-performance data centers.

[0033] System self-learning optimization module: This module continuously collects historical operation and maintenance data, management strategy execution effectiveness, and training feedback information from high-performance computing data centers. It constructs a large-scale labeled dataset and iteratively optimizes the capability level assessment model, risk level calculation model, and strategy matching algorithm through reinforcement learning algorithms. It dynamically adjusts the calculation weights of Cs and Rt, the mapping rules of the two-dimensional interlocking matrix, and the correlation of training courses, gradually improving the system's adaptability to complex operation and maintenance scenarios, the accuracy of risk identification, and the effectiveness of management strategies to meet the ever-growing operation and maintenance needs of high-performance computing data centers.

[0034] Example 1: Application of daily operation and maintenance scenarios in large-scale Internet data centers This embodiment is applied to a large-scale Internet data center, which deploys core facilities such as server clusters, cooling systems, and power supply equipment. Daily operation and maintenance involves various tasks such as equipment inspection, fault diagnosis, and parameter optimization. The center has 30 maintenance personnel and requires dynamic personnel assessment, precise task matching, and full-process risk management. The system operation flow is as follows: I. Execution of Core Processes and Key Steps System Initialization and Parameter Configuration: Maintenance personnel enter basic information about themselves and configure system operating parameters through the management workstation of the human-machine interface module. Capability levels Cs1-5 are set to correspond to scores of 0-20, 21-40, 41-60, 61-80, and 81-100 points, respectively; risk levels Rt1-5 correspond to low to extremely high risk. Communication connections and self-tests of each module are completed to ensure the normal operation of data acquisition equipment, transmission links, and storage systems. A data acquisition and transmission frequency of twice per second is set.

[0035] Operations and Maintenance Personnel Data Collection: The operations and maintenance personnel data collection module is deployed according to a preset layout. Biometric technology is used for qualification certification collection. The operation log capture tool captures the operation records of the operations and maintenance terminals in real time. Wearable fatigue sensors monitor continuous operation and maintenance duration and biometric characteristics. The compliance detection terminal verifies the HSE training certificate and annual medical examination report. The collected data includes: a certain operations and maintenance personnel's qualification certificate is complete and the qualification compliance standardization score is A=90 points; no violations in the past 3 months and the operation standardization score is B=85 points; successfully handled 5 faults and the fault handling ability standardization score is C=80 points; passed the annual medical examination and the health suitability standardization score is D=95 points; the HSE training is fully met and the training compliance rate standardization score is E=90 points. The real-time operation log shows that the continuous operation and maintenance lasted for 2 hours and the mental state was good.

[0036] Operation and maintenance task information collection: The operation and maintenance task information collection module obtains task information through the Internet of Things sensor network. This task is to optimize the parameters of the server cluster. The target equipment has a high importance level, the operation and maintenance complexity is moderate, the working environment temperature and humidity is 24℃ and 45%, the power supply stability is good, the equipment has been in operation for 2 years, the core HSE risk point at the work site is electrostatic protection, the task time window is 14:00-18:00 on the same day, the fault tolerance threshold is low, and 3 core servers are associated.

[0037] Capability and Risk Level Calculation: The operations and maintenance personnel capability level assessment module uses machine learning algorithms to remove outlier data and calculates the capability score using the formula Cs=w1A+w2B+w3C+w4D+w5E, where w1=0.2, w2=0.25, w3=0.3, w4=0.15, and w5=0.1. Substituting the data, we get Cs=0.2×90+0.25×85+0.3×80+0.15×95+0.1×90=18+21.25+24+14.25+9=86.5 points, corresponding to Cs=level 5. The task risk level calculation module uses the analytic hierarchy process (AHP), combining the weights of equipment importance (0.3), operational complexity (0.25), environmental stability (0.15), task urgency (0.15), and historical failure frequency (0.15), to obtain Rt=level 3.

[0038] Strategy Matching and Dynamic Control: The two-dimensional interlocking matrix and rule base module match the control policies corresponding to Cs=5 and Rt=3 levels, including allowing operation on core servers, single maintenance duration not exceeding 4 hours, no auxiliary personnel configuration required, and compatibility with high-precision maintenance tools. The dynamic control execution module issues policies. If maintenance personnel start work at 14:00, and at 16:30 the fatigue sensor detects abnormal blinking frequency, indicating continuous maintenance for 2.5 hours, the system updates Cs=4 in real time, triggering a control policy adjustment, pushing a rest reminder, extending the maintenance interval to 30 minutes, and adjusting task priorities to ensure operational accuracy.

[0039] Anomaly Handling and Training Triggering: During maintenance, a novice maintenance worker caused a brief fluctuation in server parameters due to improper operation. The system automatically recorded this anomaly in their skill profile. The targeted training trigger module matched the training with a practical course on security regulations, generating a personalized training plan including 3 hours of video tutorials, 2 virtual simulation operations, and 1 online assessment, which was then pushed to the maintenance worker's mobile device. After the training, a skill reassessment was automatically triggered, improving their operational compliance score from 65 to 80 points, and their CS level from level 3 to level 4.

[0040] Cross-regional collaboration and data security: The data storage and management module categorizes and stores all data from this operation and maintenance, supporting queries by personnel ID and task number. The data security and compliance audit module employs end-to-end encrypted transmission, records operation logs using blockchain technology, and sets permissions for administrators to view only data within their authorized scope. The system self-learning optimization module collects the data from this operation and maintenance, adjusts the Cs calculation weights, and optimizes the mapping relationship of the two-dimensional interlocking matrix.

[0041] II. Data Representation and Interpretation Table 1 compares the daily operation and maintenance performance of large-scale Internet data centers: Table 1

[0042] Table 1 clearly demonstrates the advantages of this invention in daily operation and maintenance scenarios. Traditional operation and maintenance methods rely on manual task allocation, lack a scientific evaluation system, and have a capability-risk matching accuracy rate of only 65%, with a failure rate of 12% due to human error. Training is mostly standardized courses, failing to address individual personnel weaknesses, resulting in insufficient targeting and a satisfaction rate of only 50%. Data management is fragmented and security measures are weak, with a compliance rate of only 75%. This invention, through multi-dimensional data collection and precise algorithm calculation, improves the capability level matching accuracy to 94%, and its dynamic adjustment mechanism promptly avoids risks such as fatigue operation, reducing the failure rate to 3%. Personalized training directly addresses capability weaknesses, achieving a satisfaction rate of 88%. Cross-regional collaboration and data security mechanisms ensure efficient operation and maintenance, and a 100% data compliance rate meets regulatory requirements, comprehensively improving the security, accuracy, and efficiency of daily operation and maintenance in large data centers.

[0043] Example 2: Application of high-risk task operation and maintenance scenarios in financial industry data centers This embodiment is applied to a core data center in the financial industry. This center undertakes massive transaction data storage and processing tasks. This maintenance task is a core database upgrade, which is of high risk and requires cross-regional scheduling of maintenance resources to ensure zero errors in the task. The system operation process is as follows: I. Execution of Core Processes and Key Steps System initialization and cross-regional configuration: Operations and maintenance personnel configure system parameters through the human-computer interaction module, setting Rt=5 as extremely high-risk tasks and Cs≥4 as the appropriate personnel level standard. The system networking between the local data center and the off-site backup data center is completed, enabling real-time data sharing through cloud communication technology. Biometric identification and hierarchical access control mechanisms are configured to ensure secure cross-regional data transmission and access.

[0044] Operations and Maintenance Personnel and Task Data Collection: The operations and maintenance personnel data collection module synchronously collects data from local and remote operations and maintenance personnel. For a senior operations and maintenance personnel located remotely, the standardized scores are as follows: Compliance score A=95, Operational Standardization score B=90, Fault Handling Ability score C=92, Health Compatibility score D=88, and Training Compliance Rate score E=93. The real-time status is one hour of continuous operations and maintenance without signs of fatigue. The operations and maintenance task information collection module obtains core database upgrade task information. The equipment is of extremely high importance, the operation is extremely complex, the operating environment temperature and humidity are 23℃ and 42%, the power supply stability is good, the historical failure frequency is low, the task time window is 00:00-04:00 the next day, the fault tolerance threshold is zero, and all transaction servers are associated.

[0045] Level Calculation and Strategy Matching: The senior operations and maintenance personnel's capability score is calculated using a formula: w1=0.2, w2=0.25, w3=0.3, w4=0.15, w5=0.1. Cs=0.2×95+0.25×90+0.3×92+0.15×88+0.1×93=19+22.5+27.6+13.2+9.3=91.6 points, corresponding to Cs=level 5. The task risk level calculation module obtains Rt=5 through multi-factor weighted calculation. A two-dimensional interlocking matrix is ​​used to match the control strategy: allowing operation of the core database, single operation and maintenance duration not exceeding 4 hours, dispatching one additional technical backbone for assistance, using dedicated operation and maintenance tools, and locking non-operation and maintenance areas. Critical Value Strategy Matching and Dynamic Adjustment: Because Cs=5 and Rt=5 are in a high-level critical state, the intelligent matching module of operation and maintenance control strategy uses a formula... The matching degree was calculated, where α=0.6, β=0.4, γ=1, and M=(91.6×0.6+5×0.4) / 1=54.96+2=56.96 points. The system selected the top 3 control strategies based on their scores, and the management personnel chose the optimized solution of dispatching more local technical personnel to collaborate. At 01:30, the environmental monitoring terminal detected a slight fluctuation in the power supply voltage. The system dynamically corrected the Rt level to level 5, adjusted the control strategy accordingly, activated the backup power supply, and shortened the operation interval to 15 minutes to ensure zero errors in the task.

[0046] Cross-regional collaboration and training triggering: The cross-regional operation and maintenance collaboration and control module generates a cross-regional operation and maintenance plan, clearly defining the travel arrangements, task handover process, and remote operation permission configuration for senior operation and maintenance personnel arriving locally on the same day. The operation and maintenance personnel complete the task handover according to the plan, and the database upgrade operation starts precisely at 00:00. The entire process proceeds without any abnormal operations, and the task is successfully completed at 03:30. The system's self-learning optimization module collects data from this high-risk task and optimizes the weighting of risk level calculation factors. For a local Cs=3 level operation and maintenance personnel, a basic skills enhancement course is matched, automatically triggering a capability reassessment process.

[0047] Data storage and compliance audit: The data storage and management module categorizes and stores cross-regional operation and maintenance data, level assessment results, policy execution records, and other full-process information. The data security and compliance audit module regularly generates compliance audit reports, in which key indicators such as the compliance rate of operation and maintenance personnel qualifications, training compliance rate, risk control execution rate, and anomaly handling timeliness rate are all 100%, fully meeting the strict compliance and regulatory requirements of the financial industry.

[0048] II. Data Representation and Interpretation Table 2 compares the operational effectiveness of high-risk tasks in the financial industry: Table 2

[0049] Table 2 data highlights the advantages of this invention in adapting to high-risk task operation and maintenance scenarios. High-risk tasks in the financial industry have extremely high requirements for security and accuracy. Traditional operation and maintenance methods lack cross-regional rapid scheduling mechanisms, resulting in response times as long as 24 hours. Strategy adaptation relies on manual experience with an accuracy of only 60%, and delayed anomaly handling leads to a task success rate of only 78%. Inadequate data management reduces compliance audit pass rates to 80%, making it difficult to meet zero-error requirements. This invention achieves rapid scheduling within 4 hours through a cross-regional collaborative control module. A two-dimensional interlocking matrix and intelligent matching algorithm ensure strategy adaptation accuracy of 95%. Dynamic control and real-time monitoring mechanisms ensure a timely anomaly handling rate of 98%, increasing the success rate of high-risk tasks to 99%. End-to-end data encryption and compliance audit mechanisms guarantee a 100% pass rate. The system comprehensively solves the pain points of slow scheduling, poor adaptation, and high risk in traditional high-risk operation and maintenance, providing strong technical support for the secure execution of core tasks in financial data centers.

[0050] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A high-computing-power data center environment awareness and coordinated operation and maintenance system, characterized in that, include: Operation and maintenance personnel data collection module: It is equipped with a qualification certification collection unit, an operation behavior monitoring unit, a health status detection unit and a training record integration unit. Through biometrics and sensor technology, it collects qualification, operation, health and training data, and after standardized conversion and verification, it transmits the data to the core processing module twice per second. Operation and maintenance task information collection module: Through the Internet of Things sensor network and the device management system interface, it obtains information on data center operation and maintenance task types, equipment importance, and environmental parameters, and synchronously collects time window constraints to form a structured task information package for transmission to the core processing module. Operations and maintenance personnel capability level assessment module: Deeply processes the data collected by operations and maintenance personnel, removes anomalies through machine learning algorithms, quantifies the score from five dimensions of qualification and compliance, and dynamically adjusts the weights in combination with real-time status to generate an operations and maintenance personnel capability status level Cs from 1 to 5. Task risk level calculation module: The analytic hierarchy process is used to analyze the operation and maintenance task information. It integrates factors such as equipment importance, operation complexity, environmental stability, task urgency, and historical failure frequency, quantifies the weight of each factor and performs weighted calculation to generate the task risk level Rt. The level is divided into 1 to 5 levels, corresponding to risks from low to high. Two-dimensional interlock matrix and rule base module: Constructs a two-dimensional interlock matrix Cs and Rt, and presets control policies corresponding to different combinations of capability levels and risk levels, including access permissions for operation and maintenance personnel, target device matching type, single operation and maintenance duration limit, number of auxiliary operation and maintenance personnel, and operation and maintenance tool adaptation standards. The rule base supports dynamic updating of mapping relationships based on actual operation and maintenance results. Dynamic control and execution module: Receives real-time data on the capability level and task risk level of operation and maintenance personnel, matches and executes control strategies according to a two-dimensional interlock matrix, continuously monitors changes in Cs and Rt during task execution, and immediately adjusts the operation and maintenance plan when either dimension exceeds the preset threshold. Targeted training trigger module: The system automatically writes back the operation and maintenance personnel's competency profiles with anomalies such as violations and improper fault handling, matches them with HSE training courses to generate personalized plans and push them out, and triggers a competency reassessment after the training is completed to update the Cs level; Human-Computer Interaction Module: The system configuration management workstation and personnel mobile terminal interaction interface, which displays Cs level, Rt level and control policy information in real time, supports managers to adjust rules and set thresholds, and allows maintenance personnel to view files, receive notifications and feedback.

2. The high-computing-power data center environment perception and linkage operation and maintenance system according to claim 1, characterized in that, Also includes: The module for calculating the capability status of operations and maintenance personnel quantifies capability status using a weighted summation algorithm. It generates a capability status score for operations and maintenance personnel using the formula Cs=w1A+w2B+w3C+w4D+w5E, where Cs is the capability status score, w1 is the qualification compliance weight coefficient, A is the qualification compliance standardization score, w2 is the operational standardization weight coefficient, B is the operational standardization standardization score, w3 is the fault handling capability weight coefficient, C is the fault handling capability standardization score, w4 is the health adaptability weight coefficient, D is the health adaptability standardization score, w5 is the training pass rate weight coefficient, and E is the training pass rate standardization score. Each weight coefficient is optimized through training using historical operations and maintenance data from a high-performance data center and satisfies w1+w2+w3+w4+w5=1. A higher score corresponds to a higher capability level, thus achieving a quantitative assessment of the capability status of operations and maintenance personnel.

3. The high-computing-power data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: The dynamic risk warning module for operation and maintenance tasks dynamically adjusts the risk level by monitoring environmental changes, equipment status fluctuations, and operation and maintenance progress in real time during task execution. When the ambient temperature and humidity exceed the safe range, equipment alarms occur suddenly, or operation and maintenance delays exceed the preset threshold, the risk level upgrade process is automatically initiated, and warning information is pushed to the management terminal and operation and maintenance personnel terminals simultaneously. The control strategy is adjusted in combination with the current CS level, and measures such as adding linkage of environmental control equipment and dispatching more technical support personnel to shorten the operation and maintenance interval are implemented.

4. The high-computing-power data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: Operation and maintenance management strategy matching module: Based on a two-dimensional interlocking matrix and rule base, it optimizes the strategy mapping accuracy through a fuzzy matching algorithm. When Cs and Rt are at the level threshold, it uses a formula... Calculate the strategy matching score, where M is the strategy matching score; α is the capability level influence coefficient, β is the risk level influence coefficient, and γ is the normalization coefficient. An analysis report of the expected effects of each strategy is also provided.

5. The high-computing-power data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: The fatigue monitoring and intervention module for maintenance personnel analyzes fatigue levels by measuring continuous maintenance duration, operation response speed, frequency of erroneous operations, and blink frequency (biological characteristics). It sets three-level fatigue warning thresholds and records fatigue data to an individual's competency profile, serving as an important basis for health suitability scoring.

6. The high-computing-power data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: High-performance computing equipment operation and maintenance adaptation and optimization module: Establish a database of equipment operation and maintenance characteristics, store operation and maintenance specifications, risk point distribution, and personnel capability requirements for different equipment. When the task risk level Rt=5, automatically match the equipment-specific operation and maintenance tool list, operation step guide, and emergency response plan from the database, configure equipment operation permissions in conjunction with the operation and maintenance personnel's Cs level, and simultaneously link with the equipment management system to lock non-operation and maintenance areas.

7. The high-performance data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: HSE training personalized customization module: Based on abnormal data, a course-abnormal type association mapping library is built. Corresponding courses are matched according to violations and improper handling of faults. Video teaching mode is adopted, training assessment data is recorded, and the pass rate is updated and a capability reassessment is triggered after passing the test, forming a closed loop.

8. The high-performance data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: Cross-regional operation and maintenance collaborative control module: Real-time sharing of data, task information and control strategies of operation and maintenance personnel in different regions is realized through cloud communication technology. When a high-risk task Rt=5 occurs in a single data center and there are no matching high-capability level Cs≥4 operation and maintenance personnel locally, the module automatically dispatches suitable personnel from other regions to generate a cross-regional operation and maintenance plan, and synchronously updates the records of data storage modules in each region to realize the optimized configuration and collaborative control of operation and maintenance resources in each data center.

9. The high-computing-power data center environment awareness and linkage operation and maintenance system according to claim 1, characterized in that, Also includes: Data security and compliance audit module: End-to-end encryption is used to protect the transmission and storage of sensitive information. Blockchain is used to prevent tampering and ensure traceability of operation logs. Hierarchical access permissions are set, and compliance audit reports including core indicators of qualification compliance rate are generated regularly.

10. The high-performance data center environment awareness and coordinated operation and maintenance system according to claim 1, characterized in that, Also includes: System self-learning optimization module: Collects historical operation and maintenance data, strategy execution effects and training feedback, builds labeled datasets, realizes reinforcement learning algorithm iterative optimization evaluation, calculation and matching model, dynamically adjusts weights and rules, and improves the system's adaptability, recognition and control capabilities.

Citation Information

Patent Citations

  • Intelligent management system for wind power operation and maintenance

    CN119205083A

  • Data center infrastructure operation and maintenance management system and method

    CN119512878A

  • Water affair facility operation and maintenance detection method and system based on big data

    CN121189777A

  • Intelligent control method and system for special equipment based on Internet of Things

    CN121302341A

  • Server operation and maintenance method and electronic equipment

    CN121436967A