Data center operation and maintenance service system

By introducing an automated operation and maintenance service system into the data center, using the combination of equipment monitoring modules, data acquisition modules and central servers, the shortcomings of traditional operation and maintenance management in large-scale and complex environments are solved, efficient operation and maintenance management and resource allocation are achieved, and the overall performance and efficiency of the data center are improved.

CN120011173APending Publication Date: 2025-05-16CHINA SOUTHERN POWER GRID INTERNET SERVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510088913.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Traditional data center operation and maintenance management relies on manual monitoring and empirical judgment, and it is difficult to meet the needs of modern data centers for efficient operation and maintenance, especially in large-scale and complex environments.

Method used

Provide a data center operation and maintenance service system, including equipment monitoring module, data acquisition module and central server. The system uses automated data monitoring, extract performance indicator data, analyzes data change trends and correlation modes, determines the best operation and maintenance cycle, and calls the trained workload prediction model to predict the equipment's workload data, determines the computing resource allocation plan, and performs resource scheduling.

Benefits of technology

Through automated operation and maintenance management, the impact of operation and maintenance on the operation stability of data centers is reduced, the accuracy and applicability of resource allocation are improved, and the performance and efficiency of overall operation of data centers are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011173A_ABST
    Figure CN120011173A_ABST
Patent Text Reader

Abstract

The invention relates to a data center operation and maintenance service system. The system comprises an equipment monitoring module, a data acquisition module and a central server. The equipment monitoring module sends the extracted performance index data to the central server; the central server analyzes a data change trend and a data association mode of the performance index data, and determines an optimal operation and maintenance period of the data center based on an obtained data change trend analysis result and an obtained data association mode analysis result; the data acquisition module sends the workload data of the equipment in the optimal operation and maintenance period to the central server; and the central server predicts the workload data of the equipment according to the workload data of the equipment in the optimal operation and maintenance period and a trained workload prediction model to obtain a workload prediction result, determines a computing resource allocation scheme of the equipment based on the workload prediction result and a data association mode analysis result, and performs resource scheduling. The method is beneficial to improving the operation and maintenance efficiency of the data center.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of information technology, and in particular to a data center operation and maintenance service system. Background Art

[0002] With the advent of the digital age, data centers have become the core facilities for data storage, processing and exchange in enterprise operations. Data Center Operations (DCO) is crucial to ensure the stable and efficient operation of data centers.

[0003] Traditional data center operation and maintenance management mainly relies on manual monitoring and empirical judgment, and the operation and maintenance solutions are automatically executed based on rules and simple algorithms. However, this solution is incapable of coping with large-scale and complex data center environments, and it is difficult to meet the high-efficiency requirements of modern data centers. Summary of the invention

[0004] Based on this, it is necessary to provide a data center operation and maintenance service system that can improve the operation and maintenance efficiency of data centers in response to the above technical problems.

[0005] The present application provides a data center operation and maintenance service system, including an equipment monitoring module, a data acquisition module and a central server:

[0006] The equipment monitoring module extracts performance indicator data from the historical operation records of the data center and sends the performance indicator data to the central server;

[0007] The central server analyzes the data change trend of the performance indicator data to obtain the data change trend analysis result, analyzes the data association pattern of the performance indicator data to obtain the data association pattern analysis result, and determines the optimal operation and maintenance cycle of the data center based on the data change trend analysis result and the data association pattern analysis result;

[0008] The data acquisition module obtains the workload data of the equipment within the optimal operation and maintenance cycle and sends the workload data to the central server;

[0009] The central server uses the workload data of the equipment within the optimal operation and maintenance cycle as input, calls the trained workload prediction model, predicts the workload data of the equipment in a preset future time period, and obtains the workload prediction results. Based on the workload prediction results and the data association pattern analysis results, the computing resource allocation plan of the equipment is determined, and resources are scheduled for each device based on the computing resource allocation plan.

[0010] In one of the embodiments, the device monitoring module is further used to obtain operation data and environmental data of the device, and send the operation data and environmental data to the central server;

[0011] The central server is also used to obtain fault prediction results based on operating data, environmental data and trained fault prediction models, and to issue early warnings based on the fault prediction results.

[0012] In one of the embodiments, the data acquisition module is also used to obtain alarm information and send the alarm information to the central server;

[0013] The central server is also used to take the alarm information as input, call the trained alarm event determination model, determine the category of the alarm event, determine the processing priority of the alarm event based on the category of the alarm event and the preset processing priority evaluation criteria, and push the alarm information and generate an operation and maintenance work order for the alarm event for the alarm event whose processing priority meets the preset processing priority threshold.

[0014] In one of the embodiments, the device monitoring module is further used to obtain historical event records, device operating parameters and operation and maintenance operation logs of the device, and send the historical event records, device operating parameters and operation and maintenance operation logs to the central server;

[0015] The central server determines the fault type, as well as the fault characteristics, fault causes and maintenance plans corresponding to each fault type based on historical event records, equipment operating parameters and operation and maintenance operation logs, and builds an operation and maintenance knowledge base based on the fault characteristics, fault causes and maintenance plans corresponding to each fault type.

[0016] In one embodiment, the device monitoring module is further configured to:

[0017] In the event of a device failure, the operating data and environmental data of the failed device are sent to a central server;

[0018] The central server determines the fault type based on the received operation data and environmental data, selects the maintenance plan that matches the fault type from the operation and maintenance knowledge base, and pushes the maintenance plan.

[0019] In one of the embodiments, the data acquisition module is further used to obtain power consumption data, cooling demand data, and equipment workload data of the data center, and send the power consumption data, cooling demand data, and equipment workload data to the central server;

[0020] The central server is also used to analyze the energy usage trend of the data center based on the power consumption data, cooling demand data and equipment workload data, obtain energy usage trend analysis results, and adjust the working mode of the cooling system and power supply system of the data center based on the energy usage trend analysis results.

[0021] In one of the embodiments, the data acquisition module is further used to obtain power consumption data and operating parameters of the device, and send the power consumption data and operating parameters to the central server;

[0022] The central server is also used to use power consumption data and operating parameters as input, call the trained fault point prediction model, predict the potential fault points of the equipment, obtain the fault point prediction results, and issue early warnings based on the fault point prediction results.

[0023] In one embodiment, the equipment monitoring module is also used to monitor the equipment's operating data, maintenance cycle and historical fault data, integrate the equipment's operating data, maintenance cycle and historical fault data, obtain the equipment file of the equipment, and generate an equipment upgrade plan or equipment replacement plan based on the equipment file.

[0024] In one of the embodiments, the equipment monitoring module is also used to determine the access rights of the access control system of the data center based on the equipment profile, identify whether there are abnormal access events based on the historical access records and access rights of the access control system, obtain the equipment status data of the equipment, and identify whether there are abnormal equipment operation events based on the equipment status data and historical normal operation modes. When abnormal access events and / or abnormal equipment operation events are identified, alarm information is pushed through an encrypted channel.

[0025] In one of the embodiments, the system also includes a monitoring platform for visually displaying the operating data, environmental data and alarm information of the equipment.

[0026] On the one hand, compared with the traditional data center operation and maintenance management method that relies on manual monitoring and experience judgment, the above-mentioned data center operation and maintenance system performs automated data monitoring through the data center operation and maintenance service system, extracts performance indicator data, analyzes the data change trend and data correlation pattern of the performance indicator data, and determines the optimal operation and maintenance cycle, which is conducive to the operation and maintenance management of the data center based on the optimal operation and maintenance cycle, so as to reduce the impact of operation and maintenance management on the operational stability of the data center; on the other hand, the workload data of the equipment within the optimal operation and maintenance cycle is used as input, and the trained workload prediction model is called to obtain accurate workload prediction results. According to the workload prediction results and the data analysis results of the performance indicator data, the computing resource allocation plan is determined, and resource scheduling is performed according to the computing resource allocation plan, which can ensure that the computing resource allocation can not only meet daily needs but also flexibly respond to sudden traffic, improve the accuracy and applicability of resource configuration, and thus help improve the performance and efficiency of the overall operation of the data center. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0028] Figure 1 A structural block diagram of a data center operation and maintenance service system in one embodiment;

[0029] Figure 2 It is a structural block diagram of a data center operation and maintenance service system in another embodiment;

[0030] Figure 3 It is a structural block diagram of a data center operation and maintenance service system in yet another embodiment;

[0031] Figure 4 A structural block diagram of a data center operation and maintenance service system in yet another embodiment;

[0032] Figure 5 This is a structural block diagram of a data center operation and maintenance service system in yet another embodiment. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0034] In an exemplary embodiment, Figure 1 As shown, a data center operation and maintenance service system is provided, including a device monitoring module 100, a data acquisition module 200 and a central server 300. Among them:

[0035] The equipment monitoring module 100 extracts performance indicator data from the historical operation records of the data center and sends the performance indicator data to the central server 300 .

[0036] Among them, the performance indicator data is used to measure the operating efficiency and service quality of the data center. The performance indicator data may include CPU utilization, memory usage, network throughput, disk I / O rate, etc. The device monitoring module may refer to a module that performs real-time monitoring and management of various devices in the data center. The devices in the data center mainly include servers, network devices, storage devices, etc. The historical operation record may include log information of operation events that occurred in the data center.

[0037] In practical applications, the device monitoring module can obtain the historical operation records of each device in the data center, extract performance indicator data such as CPU utilization, memory usage, network throughput, and disk I / O rate from the historical operation records, and send the performance indicator data to the central server. Before sending the performance indicator data to the central server, the performance indicator data can also be preprocessed. Specifically, data cleaning processing such as removing outliers and missing values ​​is performed to ensure the integrity and accuracy of the data, and data of different units or scales are converted to the same standard for subsequent analysis.

[0038] The central server 300 analyzes the data change trend of the performance indicator data to obtain the data change trend analysis result, analyzes the data association pattern of the performance indicator data to obtain the data association pattern analysis result, and determines the optimal operation and maintenance cycle of the data center based on the data change trend analysis result and the data association pattern analysis result.

[0039] Among them, the optimal operation and maintenance cycle represents the ideal time frame for maintaining and optimizing the data center.

[0040] In actual applications, the central server receives performance indicator data, which can be analyzed by time series analysis algorithms (such as ARIMA (Autoregressive Integrated Moving Average Model) or Holt-Winters method) to analyze the data change trend of performance indicator data between weekdays and weekends, daytime and nighttime, capture the periodicity and seasonal fluctuations of performance indicator data, and obtain data change trend analysis results. It can also predict the trend of performance indicator data in the future by establishing a data model to further analyze the data change trend. Apply rule association algorithms (such as Apriori (priori algorithm) algorithm or FP-Growth (Frequent Pattern Growth, FP-Growth, frequent pattern growth) algorithm) to analyze the potential association patterns between different performance indicator data and obtain data association pattern analysis results. For example, whether high CPU utilization is often accompanied by specific network traffic peaks. Performance indicator data can also be analyzed by clustering analysis to group similar operation records to discover potential data association patterns and obtain data association pattern analysis results.

[0041] According to the data change trend analysis results and the data association pattern analysis results, the optimal operation and maintenance cycle of the data center can be determined by first determining the business peak time period of the data center (such as 9:30 am to 3:00 pm on weekdays) based on the data change trend analysis results, and determining the non-business peak time period as the preliminary optimal operation and maintenance cycle. According to the data association pattern analysis results, the analysis takes into account factors such as equipment load, maintenance cycle of the data center, and operating status of the equipment (whether the equipment is working normally) to determine the optimal operation and maintenance cycle.

[0042] For example, for a financial institution's data center, the daytime on weekdays is usually the peak business period (such as the securities trading period), and large-scale maintenance operations are not suitable during this period. By analyzing the performance indicator data of the past few years, the data change trend analysis results and data association pattern analysis results indicate that the resource utilization rate of the data center is low during a certain period of time, then the period can be determined as the optimal operation and maintenance cycle, and preventive maintenance work can be arranged during the optimal operation and maintenance cycle to reduce the impact of operation and maintenance on the continuity of data center services.

[0043] In other embodiments, the central server is also used to enable the system to dynamically adjust the optimal operation and maintenance cycle according to the actual operation of the data center through a self-learning algorithm (such as Q-learning or deep reinforcement learning in reinforcement learning).

[0044] The data collection module 200 obtains the workload data of the equipment within the optimal operation and maintenance cycle, and sends the workload data to the central server 300 .

[0045] In actual applications, after determining the optimal operation and maintenance cycle, the data acquisition module obtains the workload data of the equipment in the data center during the optimal operation and maintenance cycle, and sends the workload data to the central server.

[0046] The central server 300 takes the workload data of the equipment within the optimal operation and maintenance cycle as input, calls the trained workload prediction model, predicts the workload data of the equipment within a preset future time period, obtains the workload prediction result, determines the computing resource allocation plan of the equipment based on the workload prediction result and the data association pattern analysis result, and schedules resources for each device based on the computing resource allocation plan.

[0047] Among them, the computing resource allocation plan is used to reasonably arrange the computing resources of servers, storage devices and other basic equipment in the data center to ensure that these devices run efficiently under the expected workload.

[0048] In practical applications, such as Figure 2As shown, the central server also includes a workload prediction module 310, which is used to use the workload data of the equipment within the optimal operation and maintenance cycle as input, call the trained workload prediction model, predict the workload data of the equipment in the preset future time period, and obtain the workload prediction result. Specifically, the workload prediction model can be pre-built according to the machine learning model (such as random forest, support vector machine, etc.), obtain the historical workload data of the equipment in the data center in the historical time period, use the historical workload data as input, call the initial workload prediction model, iterate the training until the model converges, and obtain the trained workload prediction model. Use the workload data of the equipment within the optimal operation and maintenance cycle as input, call the trained workload prediction model, predict the workload data of the equipment in the future period (such as within one month), and obtain the workload prediction result, which is used to consider the business peaks and flat peaks that may exist in different time periods in the future, as well as seasonal or periodic changes.

[0049] Based on the workload prediction results and the data association pattern analysis results, determine the computing resource allocation plan for the devices in the data center. This can be to determine the business peak and off-peak periods based on the data association pattern analysis results. According to the predicted workload, increase computing resources during peak periods (such as adding more virtual machine instances or containers), reduce computing resources during off-peak periods to reduce costs, and dynamically adjust computing resources according to the workload of each device to avoid single-point overload and determine the computing resource allocation plan. Execute corresponding resource scheduling commands according to the computing resource allocation plan, such as starting or stopping containers, migrating virtual machines, adjusting network bandwidth, etc.

[0050] Continuing with the above example, considering that the financial market has obvious daily volatility, it is expected that there will be higher computing demand in the morning before the market opens and after the market closes. Through the workload prediction results and data association pattern analysis results, sufficient computing resources can be predicted and prepared in advance to cope with the upcoming workload peak, and idle resources can be released during the trough period to save costs.

[0051] In other embodiments, after determining the computing resource allocation plan, the equipment monitoring module is also used to monitor performance indicators during the operation and maintenance cycle, obtain performance indicator data, and send the performance indicator data to the central server. The central server receives the performance indicator data and determines whether there are resource processes or areas with insufficient resources in the data center based on the performance indicator data and the reinforcement learning model, and adjusts the computing resource allocation plan accordingly, such as increasing or decreasing the number of instances of a specific service, changing the configuration of the virtual machine, etc., to take into account daily processing needs and respond to emergencies. Specifically, the optimal solution for the computing resource allocation plan is found through the trial and error learning method of the reinforcement learning model, and the execution effect of the plan is evaluated after each decision to obtain an evaluation result, and the internal parameters are updated according to the evaluation result.

[0052] In other embodiments, the central server is also used to generate a computing resource usage report based on the computing resource usage of each device after resource scheduling, so as to audit whether the computing resource allocation plan is effective and whether the resource scheduling complies with established security standards and service level agreements.

[0053] In other embodiments, the central server is also used to automatically optimize the computing resource configuration plan through genetic algorithms or other heuristic search strategies according to the optimal operation and maintenance cycle and real-time demand changes.

[0054] In the above-mentioned data center operation and maintenance service system, on the one hand, compared with the traditional data center operation and maintenance management method that relies on manual monitoring and experience judgment, the data center operation and maintenance service system performs automatic data monitoring, extracts performance indicator data, analyzes the data change trend and data correlation pattern of the performance indicator data, and determines the optimal operation and maintenance cycle, which is conducive to the operation and maintenance management of the data center based on the optimal operation and maintenance cycle, so as to reduce the impact of operation and maintenance management on the operational stability of the data center; on the other hand, the workload data of the equipment within the optimal operation and maintenance cycle is used as input, and the trained workload prediction model is called to obtain accurate workload prediction results. According to the workload prediction results and the data analysis results of the performance indicator data, the computing resource allocation plan is determined, and resource scheduling is performed according to the computing resource allocation plan, which can ensure that the computing resource allocation can not only meet daily needs but also flexibly respond to sudden traffic, thereby improving the accuracy and applicability of resource configuration, which is conducive to improving the performance and efficiency of the overall operation of the data center.

[0055] In other embodiments, the central server is also used to adjust the data collection frequency of the device's sensors based on the workload prediction results, and to monitor the devices in advance for those devices whose loads are expected to increase significantly in the future; otherwise, the collection frequency is appropriately reduced to save energy.

[0056] In an exemplary embodiment, the equipment monitoring module 100 is also used to obtain operation data and environmental data of the equipment, and send the operation data and environmental data to the central server 300 .

[0057] The operation data includes but is not limited to current, power, vibration frequency, etc. The environmental data includes but is not limited to temperature and humidity, etc.

[0058] In practical applications, the equipment monitoring module may obtain the equipment's operating data (such as current, voltage and vibration data) and environmental data (such as temperature and humidity) from various sensors or other data sources, and perform data preprocessing on the operating data and environmental data, including but not limited to cleaning abnormal values, filling missing values, standardization, normalization, etc. The operating data and environmental data are sent to the central server.

[0059] The central server 300 is also used to obtain fault prediction results based on operation data, environmental data and trained fault prediction models, and issue early warnings based on the fault prediction results.

[0060] In practical applications, such as Figure 3 As shown, the central server also includes a fault prediction module 320, which is used to obtain fault prediction results based on operating data, environmental data and trained fault prediction models, and to issue warnings based on the fault prediction results. Specifically, it can be based on a deep learning framework (such as LSTM, GRU and other recurrent neural networks) or a traditional time series analysis model (such as ARIMA) in advance, and combined with other machine learning algorithms (such as random forests, XGBoost, etc.), to build an initial fault prediction model, annotate each historical operating data and historical environmental data, annotate the fault type (such as hardware failure, software failure, network failure and power system abnormality, etc.) and the fault time to obtain a training sample. With the training sample as input, the initial fault prediction model is called, and the initial fault prediction model is iteratively trained until the model converges to obtain a fault prediction model. With the operating data and environmental data as input, the trained fault prediction model is called to predict possible faults and obtain a fault prediction result. According to the fault prediction result, a fault probability curve for each device in the future period is generated, so that the operation and maintenance personnel can intuitively understand which devices are most likely to fail, integrate the fault prediction results and the fault probability curve, obtain fault warning information, and push the fault warning information. The fault warning information can be pushed through a pop-up window on the monitoring platform, or it can be sent to the terminal device of the operation and maintenance personnel.

[0061] In this embodiment, a fault prediction model is introduced to predict potential faults through the operating data and environmental data of the equipment, thereby improving the accuracy of fault prediction. By identifying potential fault risks in advance and issuing early warnings, it is beneficial to take preventive measures in time to avoid service interruptions and service quality degradation.

[0062] In an exemplary embodiment, the data acquisition module 200 is also used to obtain alarm information and send the alarm information to the central server 300 .

[0063] The alarm information may include the name of the device that triggers the alarm, the name of the service that triggers the alarm, the service port, the time when the alarm occurs, and the alarm content. The alarm content may include the cause and phenomenon of the alarm.

[0064] The data acquisition module obtains the alarm information of the device through the monitoring interface deployed on the device and sends the alarm information to the central server.

[0065] The central server 300 is also used to take the alarm information as input, call the trained alarm event determination model, determine the category of the alarm event, determine the processing priority of the alarm event based on the category of the alarm event and the preset processing priority evaluation standard, and push the alarm information and generate an operation and maintenance work order for the alarm event for the alarm event whose processing priority meets the preset processing priority threshold.

[0066] Among them, the categories of alarm events can be divided into security alarm categories, performance alarm categories, availability alarm categories, configuration alarm categories and environmental alarm categories. Among them, the security alarm category can include alarm events such as unauthorized access attempts, malware detection, firewall rule violations, etc.; the performance alarm category can include alarm events such as excessive CPU utilization, memory leaks, and abnormal disk I / O rates; the availability alarm category can include alarm events such as service interruption, hardware failure, and network connection loss; the configuration alarm category can include alarm events such as configuration file changes and permission changes; the environmental alarm category can include alarm events such as temperature exceeding the standard and humidity abnormality. The processing priority of alarm events can include but is not limited to four levels: emergency, important, general, and prompt, which respectively indicate that the system has a serious failure, an important problem, a general problem, or a minor problem.

[0067] In practical applications, such as Figure 4 As shown, the central server also includes an alarm event determination module 330, which is used to use the alarm information as input, call the trained alarm event determination model, and determine the category of the alarm event. Specifically, it can be pre-built based on other classification models such as support vector machines and decision trees to distinguish different categories of alarm events, receive historical alarm information sent by the data acquisition module, perform data annotation on each historical alarm information, annotate the category of the alarm event, and obtain a training set. Using the training set as input, call the initial alarm event determination model, iteratively train the initial alarm event determination model until the model converges, and obtain the trained alarm event determination model. The alarm event determination module 320 uses the alarm information as input, calls the trained alarm event determination model, classifies the alarm event corresponding to the alarm information, and determines the category of the alarm event.

[0068] The central server determines the impact scope, severity, and urgency of the alarm event based on the alarm event category. Set the processing priority evaluation standard in advance based on the relationship between the impact scope, severity, and urgency corresponding to the alarm event category, and determine the processing priority of the alarm event based on the comprehensive information. Specifically, if the impact scope of the alarm event involves critical business, its processing priority is determined to be urgent; if the impact scope of the alarm event does not involve critical business, but its urgency indicates that it may escalate to a more serious event if not handled in time, its processing priority is determined to be important; if its impact scope is large, but the severity and urgency are low, its processing priority is determined to be general; if its impact scope is small and has almost no impact on the business, its processing priority is determined to be prompt.

[0069] Pre-set the processing priority threshold (such as the critical level), and for alarm events with a processing priority higher than or equal to the critical level, push the alarm information to the operation and maintenance personnel, and generate an operation and maintenance work order for the alarm event. The operation and maintenance work order can include the category of the alarm event, the processing priority, the device name, the alarm cause, and the recommended solution. The recommended solution can set a mapping table between the alarm event and the recommended solution based on experience, and match the recommended solution corresponding to the alarm event through the mapping table. Push the operation and maintenance work order to the operation and maintenance personnel responsible for the area where the alarm event occurs, and continuously monitor the status changes of the operation and maintenance work order, from the entire process of receiving, processing to closing, and dynamically update the status records of related equipment accordingly to ensure that all operations are traceable. For alarm events below the processing priority threshold, they can be arranged for processing during planned maintenance.

[0070] Following the above embodiment, the financial industry has strict requirements for data security. Once abnormal behavior or potential risks are detected, such as unauthorized access attempts, alarm information is triggered, and the category and processing priority of the alarm event are determined. For alarm events whose processing priority meets the preset processing priority threshold, the alarm information is pushed. If the processing priority is higher than the general level, the alarm information is pushed. When the alarm event involves the leakage of personal privacy information, the system will give priority to processing and transmit the alarm information to the designated security manager through encryption.

[0071] In this embodiment, the category of alarm information is determined by the alarm event judgment model, which improves the accuracy and efficiency of alarm information classification. The processing priority of the alarm event is determined according to the category of the alarm information and the preset processing priority evaluation standard. For alarm events with high processing priority, the alarm information is pushed and an operation and maintenance work order is generated, which is beneficial to improving the effectiveness of operation and maintenance and the reliability of the data center.

[0072] In an exemplary embodiment, the equipment monitoring module 100 obtains historical event records, equipment operating parameters, and operation and maintenance logs of the equipment, and sends the historical event records, equipment operating parameters, and operation and maintenance logs to the central server 300 .

[0073] Among them, historical event records may include records of various events that occurred during the life cycle of the device, including but not limited to alarm and error information, fault records, startup / shutdown device events, configuration change records, software update records, maintenance records, etc. Device operating parameters may include CPU utilization, memory utilization, disk I / O rate, network throughput, power consumption, etc. Operation and maintenance operation logs are used to record operations performed on the device by administrators or other authorized personnel, and may include login / logout records, command execution records, backup and recovery operation records, software deployment records, physical access records, etc.

[0074] In actual applications, the equipment monitoring module 100 will obtain the equipment's historical event records, equipment operating parameters, and operation and maintenance operation logs and send them to the central server.

[0075] The central server 300 determines the fault type, as well as the fault characteristics, fault causes and maintenance plans corresponding to each fault type based on historical event records, equipment operating parameters and operation and maintenance operation logs, and builds an operation and maintenance knowledge base based on the fault characteristics, fault causes and maintenance plans corresponding to each fault type.

[0076] In actual applications, the central server integrates historical event records, equipment operating parameters, and operation and maintenance logs, identifies fault types through machine learning algorithms, and summarizes fault characteristics (such as alarm information, abnormal performance indicators, specific log records, physical phenomena of equipment), fault causes (such as power failure, program errors, connection terminals, unauthorized access, etc.) and maintenance plans (such as preventive measures, diagnostic steps, maintenance operations, etc.) corresponding to the fault type based on experience. Based on the relationship between the summarized fault type and the corresponding fault characteristics, fault causes, and maintenance plans, the graph construction tool is used to build an operation and maintenance knowledge base.

[0077] In other implementations, the operation and maintenance knowledge base may also store device performance optimization suggestions and best practice operating procedures corresponding to each type of device.

[0078] In other embodiments, in order to ensure the timeliness and accuracy of the operation and maintenance knowledge base content, the system sets an automatic update mechanism. Whenever a new fault event is resolved or a better operation method is discovered, the relevant information will be promptly entered into the operation and maintenance knowledge base and all operation and maintenance personnel will be notified simultaneously.

[0079] In other embodiments, the central server is also used to generate operation and maintenance training courseware based on the operation and maintenance knowledge base, and the operation and maintenance training courseware may include the best practice operation procedures and fault diagnosis of the equipment. It is also used to push operation and maintenance training courseware that matches the recent fault types to operation and maintenance trainees for learning. It is also used to customize personalized training plans for new operation and maintenance personnel based on their job requirements and personal skill levels in combination with the operation and maintenance knowledge base, and push the training plans to the corresponding operation and maintenance personnel. The content of the training plan covers multiple levels from basic equipment operation instructions to complex fault diagnosis techniques, aiming to comprehensively improve the professional quality and technical capabilities of the trainees.

[0080] In this embodiment, by building an operation and maintenance knowledge base, common fault characteristics and repair steps are automatically summarized so that operation and maintenance personnel can quickly obtain effective solutions, which is beneficial to improving the operation and maintenance efficiency and reliability of the data center.

[0081] In an exemplary embodiment, the equipment monitoring module 100 is also used to send the operating data and environmental data of the failed equipment to the central server 300 when the equipment fails.

[0082] In actual applications, the equipment monitoring module can obtain the operating data and environmental data of the equipment, and when the operating data or environmental data is abnormal (such as exceeding a preset safety threshold), determine that the equipment has failed, and send the operating data and environmental data of the failed equipment to the central server.

[0083] The central server 300 determines the fault type based on the received operation data and environment data, selects a maintenance plan that matches the fault type from the operation and maintenance knowledge base, and pushes the maintenance plan.

[0084] In actual applications, the central server uses the received operation data and environmental data as input, calls the fault prediction model, identifies the fault type, selects the maintenance plan that matches the fault type from the operation and maintenance knowledge base, and pushes the maintenance plan to the terminal device of the operation and maintenance personnel. If conditions permit, some simple maintenance actions (such as restarting services, adjusting configuration parameters, etc.) can be directly issued by the central server to the equipment monitoring module for automatic completion to speed up problem solving.

[0085] In this embodiment, when a fault occurs, a maintenance plan matching the fault type is screened out through the operation and maintenance knowledge base, which is beneficial to improving the operation and maintenance efficiency of the data center.

[0086] In an exemplary embodiment, the data collection module 200 is also used to obtain power consumption data, cooling demand data and equipment workload data of the data center, and send the power consumption data, cooling demand data and equipment workload data to the central server 300.

[0087] The power consumption data may include energy consumption data of the equipment. The cooling demand data may include total cooling capacity.

[0088] In actual applications, the data acquisition module obtains the power consumption data of each device through sensors deployed on the device. It can also obtain the total cooling capacity through the cooling device operation log. It obtains the workload data of the device and sends the power consumption data, cooling demand data and workload data of the device to the central server.

[0089] The central server 300 is also used to analyze the energy usage trend of the data center according to the power consumption data, cooling demand data and equipment workload data, obtain energy usage trend analysis results, and adjust the working mode of the cooling system and power supply system of the data center based on the energy usage trend analysis results.

[0090] In practical applications, the central server can analyze the changing trends of power consumption data, cooling demand data and equipment workload data over time through time series, determine the power consumption and cooling demand corresponding to the business peak period and the business off-peak period, and obtain the energy usage trend analysis results. Based on the energy usage trend analysis results, the working mode of the cooling system and power supply system of the data center is adjusted. It can be that according to the power consumption and cooling demand corresponding to the business peak period and the business off-peak period, it is determined which equipment or time period has resource waste, such as overcooling and excess power supply, and the operating parameters of the cooling equipment in the cooling system of the data center are adjusted accordingly, such as compressor frequency, fan speed, etc., to achieve on-demand cooling, and the power transmission of the power supply system is adjusted through the power scheduling algorithm.

[0091] In other embodiments, the central server is also used to predict energy waste time periods or equipment within a preset future time period based on power consumption data, cooling demand data, and equipment workload data through machine learning, and generate response strategies based on equipment performance optimization suggestions from an operation and maintenance knowledge base.

[0092] In other embodiments, the system further includes a closed-loop control module, which is used to determine whether the adjustment operation will cause a fault problem based on the fault diagnosis information in the operation and maintenance knowledge base after adjusting the working mode of the cooling system and the power supply system. Feedback information after the adjustment operation is implemented is collected to evaluate the energy-saving effect, and the new information is updated back to the knowledge base to form a virtuous cycle learning mechanism.

[0093] In this embodiment, the energy usage trend of the data center, the working mode of the cooling system and the power supply system are analyzed through power consumption data, cooling demand data and equipment workload data, thereby reducing the energy consumption of the data center and reducing costs.

[0094] In an exemplary embodiment, the data acquisition module 200 is also used to obtain power consumption data and operating parameters of the device, and send the power consumption data and operating parameters to the central server 300 .

[0095] The operating parameters may include vibration data and noise data during equipment operation.

[0096] In actual applications, the data acquisition module obtains the power consumption data of the equipment in different working states (such as idle, light load, and heavy load), as well as the vibration data and noise data during operation through sensors deployed on the equipment, and sends the power consumption data, vibration data, and noise data to the central server.

[0097] The central server 300 is also used to use the power consumption data and operating parameters as input, call the trained fault point prediction model, predict the potential fault point of the equipment, obtain the fault point prediction result, and issue an early warning based on the fault point prediction result.

[0098] In practical applications, an initial fault point prediction model can be built in advance based on a deep learning neural network, and the power consumption data and operating parameters in the historical time period sent by the data acquisition module can be labeled to mark the fault point and fault time to obtain training samples. The initial fault point prediction model is called with the training samples as input, and the initial fault point prediction model is iteratively trained until the model converges to obtain the fault point prediction model. The trained fault point prediction model is called with the power consumption data and operating parameters as input to predict the potential fault points of the equipment and obtain the fault point prediction results. The fault point prediction results are pushed to the operation and maintenance personnel to prompt them to arrange inspections and repairs to prevent unplanned downtime caused by sudden failures.

[0099] In other embodiments, the central server is also used to analyze the energy consumption characteristics of the device under different working conditions based on the power consumption data of the device, and adjust the task allocation of the device according to the load balancing algorithm to avoid some devices being in a high-load state for a long time while other devices are idle.

[0100] In this embodiment, based on power consumption data and operating parameters, the potential failure points of the equipment are predicted through a failure point prediction model, which is conducive to taking preventive measures according to the potential failure points, reducing the probability of failure and improving the reliability of the data center.

[0101] In an exemplary embodiment, the equipment monitoring module 100 is also used to monitor the equipment's operating data, maintenance cycles, and historical fault data, integrate the equipment's operating data, maintenance cycles, and historical fault data, obtain the equipment's equipment profile, determine the equipment's failure rate and performance degradation rate based on the equipment profile, and generate an equipment upgrade plan or equipment replacement plan.

[0102] The maintenance cycle may refer to the frequency at which the equipment requires regular maintenance.

[0103] In practical applications, the equipment monitoring module can obtain the equipment's operating data (such as current, voltage and vibration data) from various sensors or other data sources, obtain the maintenance cycle and historical fault data from the equipment log, integrate the equipment's operating data, maintenance cycle and historical fault data for each device, and can also integrate the equipment type, model, location, whether it is an important device and other data to obtain the equipment file for equipment life cycle management. According to the historical fault data in the equipment file, the failure rate of the equipment is determined, and the performance degradation rate of the equipment is determined according to the operating data and historical fault data in the equipment file. When the performance degradation rate of the equipment is higher than the preset performance degradation rate threshold and the failure rate of the equipment is lower than the preset failure rate threshold, a device upgrade plan is generated. The equipment upgrade plan may include a hardware upgrade plan (such as increasing memory, expanding storage space, updating network cards, etc.), a software upgrade plan (such as updating operating systems, drivers and applications to new versions) and a firmware upgrade plan. When the failure rate of the equipment is higher than the preset failure rate threshold, the remaining service life of the equipment is evaluated according to the performance degradation rate of the equipment, and an equipment replacement plan is generated. The equipment replacement plan may include an equipment replacement schedule and an equipment data migration plan.

[0104] In this embodiment, the equipment profile is obtained by integrating the equipment's operating data, maintenance cycle, and historical fault data. Based on the equipment profile, an equipment upgrade plan or equipment replacement plan is generated, which can automatically monitor the equipment's service life and help improve operation and maintenance efficiency.

[0105] In an exemplary embodiment, the device monitoring module 100 is also used to determine the access rights of the access control system of the data center based on the device profile, identify whether there are abnormal access events based on the historical access records and access rights of the access control system, obtain the device status data of the device, and identify whether there are abnormal device operation events based on the device status data and historical normal operation modes. When abnormal access events and / or abnormal device operation events are identified, an alarm message is pushed through an encrypted channel.

[0106] Among them, abnormal access events may include unauthorized access. Abnormal device operation events may include unexpected restarts or configuration changes. Device status data may include device performance indicator data and operation logs.

[0107] In actual applications, the device monitoring module determines the access rights of users (such as administrators, engineers, visitors, etc.) to the device based on the device type, location, and whether it is an important device in the device profile, including the time period, specific area or device allowed to access. Obtain the historical access records of the access control system. The historical access records may include the entry and exit events of the access control system, including timestamps, personnel identities, access areas, etc. Unauthorized access events are identified based on historical access records and access rights. Regularly obtain the performance indicator data and operation logs of the device, compare the device status data with the historical normal operation mode, and identify unexpected restarts or configuration changes. In the case of identifying abnormal access events and / or abnormal device operation events, generate alarm information, which can be transmitted to the management personnel through a secure communication protocol. The secure communication protocol can include strong encryption protocols such as SSL / TLS. After pushing the alarm information, the emergency response process can also be initiated, including cutting off the network connection, isolating the infected host, pushing the alarm information to the emergency response team, starting the system backup, etc.

[0108] In this embodiment, based on the equipment file, the access rights of the access control system are determined, the equipment status data is obtained, and it is detected whether there are abnormal access events and / or abnormal equipment operation events. Alarm information is pushed through an encrypted channel, which is conducive to quickly locking the area where the abnormal event occurred and initiating emergency plans, which is conducive to improving the security and reliability of the data center.

[0109] In other embodiments, the central server is also used to adjust the computing resource allocation plan and operation and maintenance strategy of the device through a machine learning algorithm based on abnormal operation events of the device and device profiles.

[0110] In an exemplary embodiment, Figure 5 As shown, the system also includes a monitoring platform 400 for visually displaying the operating data, environmental data and alarm information of the equipment.

[0111] In actual applications, the equipment monitoring module may send the acquired equipment operation data, environmental data and alarm information to the monitoring platform, and the monitoring platform may visualize the received data in a multi-dimensional monitoring dashboard, and the display content may include equipment operation status, alarm information, maintenance plan corresponding to the alarm information, equipment workload, environmental data, etc. It is also used to generate corresponding data reports in response to the display items selected by the monitoring personnel on the multi-dimensional monitoring dashboard.

[0112] In other embodiments, the data acquisition module is also used to send the acquired power consumption data of the device to the monitoring platform. The monitoring platform generates an energy consumption level curve chart based on the power consumption data of the device, and visualizes the energy consumption level curve chart so that monitoring personnel can understand the energy consumption level of the device in different time periods.

[0113] In this embodiment, the real-time data display of the monitoring platform is helpful to monitor the operation status of the data center, discover possible precursors of failures in time, and improve prevention efficiency.

[0114] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0115] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be construed as limiting the scope of the present application. It should be noted that, for a person of ordinary skill in the art, several modifications and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A data center operation and maintenance service system, characterized in that: The system includes an equipment monitoring module, a data acquisition module and a central server: The equipment monitoring module extracts performance indicator data from the historical operation records of the data center and sends the performance indicator data to the central server; The central server analyzes the data change trend of the performance indicator data to obtain a data change trend analysis result, analyzes the data association pattern of the performance indicator data to obtain a data association pattern analysis result, and determines the optimal operation and maintenance cycle of the data center based on the data change trend analysis result and the data association pattern analysis result; The data acquisition module acquires workload data of the equipment within the optimal operation and maintenance period, and sends the workload data to the central server; The central server takes the workload data of the equipment within the optimal operation and maintenance cycle as input, calls the trained workload prediction model, predicts the workload data of the equipment within a preset future time period, obtains the workload prediction result, determines the computing resource allocation plan of the equipment based on the workload prediction result and the data association pattern analysis result, and schedules resources for each device based on the computing resource allocation plan.

2. The system according to claim 1, characterized in that The equipment monitoring module is also used to obtain operation data and environmental data of the equipment, and send the operation data and the environmental data to the central server; The central server is also used to obtain a fault prediction result based on the operation data, the environmental data and the trained fault prediction model, and to issue an early warning based on the fault prediction result.

3. The system according to claim 2, characterized in that The data acquisition module is also used to obtain alarm information and send the alarm information to the central server; The central server is also used to take the alarm information as input, call the trained alarm event determination model, determine the category of the alarm event, determine the processing priority of the alarm event based on the category of the alarm event and a preset processing priority evaluation standard, and for alarm events whose processing priority meets a preset processing priority threshold, push the alarm information and generate an operation and maintenance work order for the alarm event.

4. The system according to claim 3, characterized in that The equipment monitoring module is also used to obtain historical event records, equipment operating parameters and operation and maintenance operation logs of the equipment, and send the historical event records, the equipment operating parameters and the operation and maintenance operation logs to the central server; The central server determines the fault type, as well as the fault characteristics, fault causes and maintenance plans corresponding to each fault type based on the historical event records, the equipment operating parameters and the operation and maintenance operation logs, and builds an operation and maintenance knowledge base based on the fault characteristics, fault causes and maintenance plans corresponding to each fault type.

5. The system according to claim 4, characterized in that The equipment monitoring module is also used for: In the event of a device failure, the operating data and environmental data of the failed device are sent to the central server; The central server determines the fault type based on the received operation data and the environmental data, selects a maintenance plan matching the fault type from the operation and maintenance knowledge base, and pushes the maintenance plan.

6. The system according to claim 5, characterized in that The data acquisition module is also used to obtain power consumption data, cooling demand data and equipment workload data of the data center, and send the power consumption data, the cooling demand data and the equipment workload data to the central server; The central server is also used to analyze the energy usage trend of the data center according to the power consumption data, the cooling demand data and the workload data of the equipment, obtain energy usage trend analysis results, and adjust the working modes of the cooling system and power supply system of the data center based on the energy usage trend analysis results.

7. The system according to claim 6, characterized in that The data acquisition module is also used to obtain power consumption data and operating parameters of the device, and send the power consumption data and the operating parameters to the central server; The central server is also used to use the power consumption data and the operating parameters as input, call the trained fault point prediction model, predict the potential fault point of the equipment, obtain the fault point prediction result, and issue an early warning based on the fault point prediction result.

8. The system according to claim 7, characterized in that The equipment monitoring module is also used to monitor the equipment's operating data, maintenance cycle and historical fault data, integrate the equipment's operating data, the maintenance cycle and the historical fault data to obtain the equipment file of the equipment, and generate an equipment upgrade plan or equipment replacement plan based on the equipment file.

9. The system according to claim 8, characterized in that The device monitoring module is also used to determine the access rights of the access control system of the data center based on the device profile, identify whether there are abnormal access events based on the historical access records of the access control system and the access rights, obtain the device status data of the device, and identify whether there are abnormal device operation events based on the device status data and historical normal operation modes. When abnormal access events and / or abnormal device operation events are identified, alarm information is pushed through an encrypted channel.

10. The system according to any one of claims 3 to 9, characterized in that: The system also includes a monitoring platform for visually displaying the operating data, environmental data and alarm information of the equipment.

Citation Information

Cited By

  • Fault detection method and system for computing power system based on artificial intelligence

    CN120872658A

  • Work allocation method and system for multiple types of units based on dynamic load

    CN121638594A