Data center energy-saving tuning control method and system based on AI large model

By applying energy-saving tuning and control methods based on AI models in data centers, the problem that the existing technology cannot perform refined energy consumption regulation is solved, and accurate prediction and optimization of data center energy consumption is achieved, and energy utilization efficiency and overall energy-saving effect are improved.

CN119987968AInactive Publication Date: 2025-05-13BEIJING ZHIKONGYUAN TECH CO LTD

Patent Information

Application Number
CN202510066599.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-16
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art cannot perform refined energy consumption regulation based on the real-time operation status of the data center and the service load, resulting in the inability to quickly make optimal energy-saving decisions in the face of large amounts of operational data and real-time changes.

Method used

The energy-saving tuning control method of data center based on AI large models is adopted. By obtaining data center information, determining the tuning target, selecting appropriate models for training and optimization, and generating AI large models. This model is used to predict data center energy consumption, analyze abnormal results, classify servers, schedule tasks, optimize server load and parameters, and realize refined energy consumption management.

Benefits of technology

Through AI large-scale models, the energy consumption of data centers can be accurately predicted and regulated, and equipment operating parameters can be adjusted in real time, dynamically adapted to changes in business loads, improved energy utilization efficiency, and improved overall energy saving effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119987968A_ABST
    Figure CN119987968A_ABST
Patent Text Reader

Abstract

The invention discloses a data center energy-saving tuning control method and system based on an AI large model, relates to the technical field of data center optimization, and solves the problems that refined energy consumption regulation and control cannot be carried out according to the real-time operation state and service load of a data center, and a large amount of operation data and real-time changes cannot be carried out. According to the method, the energy consumption of the data center is accurately predicted through an AI large model, the equipment operation parameters can be adjusted in time according to the prediction result, refined energy consumption management is realized, interaction with the data center environment can be realized in real time, the equipment operation parameters can be adjusted in real time, and the energy-saving efficiency of the data center is improved. According to the method, the service load change is dynamically adapted, the energy utilization efficiency is improved, meanwhile, the tasks to be processed are reasonably allocated to the low-energy-consumption servers through reasonable scheduling optimization of the tasks and the servers, and the overall energy-saving effect is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data center optimization technology, and specifically to a data center energy-saving tuning control method and system based on an AI large model. Background Art

[0002] With the rapid development of information technology, data centers, as core hubs for information storage and processing, are growing in size and number. Data centers consume a large amount of energy during operation, making energy costs a significant expense in data center operations.

[0003] The patent application with publication number CN116361703A discloses an energy-saving control method, device, electronic device and readable medium for a data center. The method includes: the data center monitoring system monitors and collects at least one type of energy consumption data of each device in the data center for a time period, and compares each type of energy consumption data of the device with the historical average energy consumption data to obtain a comparison result. Based on the comparison result, it is determined whether the device is in a healthy state. If the device is not in a healthy state, energy-saving control operations are performed on the device.

[0004] However, some existing energy-saving tuning systems are unable to perform fine-grained energy consumption regulation based on the real-time operating status and business load of the data center. Faced with large amounts of operating data and real-time changes, they cannot quickly make optimal energy-saving decisions. Summary of the Invention

[0005] In response to the shortcomings of the existing technology, the present invention provides a data center energy-saving tuning control method and system based on an AI large model, which solves the problem of being unable to perform fine-grained energy consumption regulation according to the real-time operating status and business load of the data center, and unable to quickly make optimal energy-saving decisions in the face of large amounts of operating data and real-time changes.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a data center energy-saving tuning control method based on an AI large model, the method specifically comprising the following steps:

[0007] Step S001: Acquire data center information and determine the optimization target. Select a suitable model based on the optimization target, and train and optimize the model to obtain an AI large model.

[0008] Step S002: Substitute the real-time data center energy consumption into the AI ​​big model for prediction to obtain predicted data, and compare the predicted data with the normal value to generate a comparison result, and the comparison result includes a normal result and an abnormal result;

[0009] Step S003: Analyze the abnormal results obtained, classify the servers according to their real-time energy consumption, and obtain server classification results. For high-energy-consuming servers, optimize the current real-time processing tasks by combining historical data to generate tasks to be processed.

[0010] Step S004: Obtaining pending tasks and low-energy servers, determining the real-time load of the low-energy servers, and performing scheduling optimization processing on the low-energy servers that meet the real-time load requirements to generate tuning control information;

[0011] Step S005 , adjusting the low-energy consumption server according to the tuning control information, and performing tuning analysis on the CPU frequency and voltage parameters of the server according to the actual workload of the adjusted server to generate secondary tuning information.

[0012] As a further solution of the present invention, the specific method of obtaining the AI ​​large model in step S001 is:

[0013] Obtain data center information, then determine the data center characteristics and energy-saving targets based on the data center information, and simultaneously obtain a learning model suitable for the data center characteristics and a learning model corresponding to the energy-saving targets, and select a common model that is suitable for both. Use this as the standard to determine the selected model, and then train the model on the preprocessed data based on the determined selected model. Optimize the model by adjusting the model parameters and hyperparameters to obtain a large AI model.

[0014] As a further solution of the present invention, the specific method of generating the comparison result in step S002 is:

[0015] The real-time energy consumption corresponding to the data center is obtained, and the obtained real-time energy consumption is substituted into the AI ​​large model for analysis to obtain predicted data. At the same time, the obtained predicted data is compared and analyzed with the normal value. If the predicted data is greater than the normal value, an abnormal result is generated. Conversely, if the predicted data is less than the normal value, a normal result is generated.

[0016] As a further solution of the present invention, the specific method of analyzing the abnormal results in step S003 is:

[0017] The server corresponding to the abnormal result is recorded as an abnormal server. At the same time, the real-time energy consumption of the abnormal server is obtained and compared with the preset energy consumption value. If the real-time energy consumption is greater than the preset energy consumption value, it is classified as a high-energy consumption server, otherwise it is classified as a low-energy consumption server;

[0018] Then obtain all high-energy-consuming servers labeled i, and i = 1, 2, ..., j, where j represents the number of high-energy-consuming servers. At the same time, obtain the real-time processing tasks of the high-energy-consuming servers, and classify them into high-load tasks and low-load tasks according to the energy consumption requirements of the real-time processing tasks, and then analyze the two.

[0019] As a further solution of the present invention, the specific method of analyzing the high-load task and the low-load task in step S003 is:

[0020] Obtain historical data of high-energy-consuming servers, obtain all processed high-load task records in the historical data, and obtain the corresponding processing times. Calculate the processing ratio values ​​corresponding to different high-load tasks and sort the processing ratio values ​​from large to small.

[0021] The sorted processing ratio values ​​are matched with the high-load tasks. If the match is successful, the corresponding high-load tasks are marked as pending tasks. If the match is unsuccessful, they are not processed.

[0022] As a further solution of the present invention, the specific method of generating the tuning control information in step S004 is:

[0023] Obtain all low-energy servers and label them as n, and n = 1, 2, ..., m, where m is the number of low-energy servers, and obtain the real-time load of the low-energy servers, and at the same time obtain the label of the pending tasks as a, and a = 1, 2, ..., b, where b represents the number of pending tasks, then sort the pending tasks a from large to small according to energy consumption requirements, and calculate the sum of the energy consumption requirements of the pending tasks and the real-time load of the low-energy servers, and at the same time compare the calculated sum with the judgment load value. If the sum is greater than the judgment load value, reallocate the pending tasks. If the sum is less than the judgment value, directly perform information scheduling optimization processing and generate tuning control information.

[0024] As a further solution of the present invention, the specific method of generating the secondary tuning information in step S005 is:

[0025] The low-energy-consuming server is adjusted using the tuning control information, and the adjusted server is recorded as the optimized server. The actual workload corresponding to the optimized server is then obtained, and the workload is divided into different levels. The CPU frequency and voltage parameters are determined according to the different levels, and the parameter adjustment interval is generated. The actual workload is then matched with the parameter adjustment interval to obtain the corresponding CPU frequency and voltage parameters, and the satisfaction of the task requirements corresponding to the optimized server is determined.

[0026] If the adjusted optimized server can meet the corresponding task requirements, the optimized server will be adjusted based on the corresponding CPU frequency and voltage parameters. Otherwise, if the task requirements are not met, the corresponding parameter adjustment range will be increased, and secondary optimization information will be generated at the same time.

[0027] A data center energy-saving tuning and control system based on an AI large model, including a data information acquisition module, a model optimization and establishment module, a server status analysis module, a tuning analysis and processing module, and a processing information output module;

[0028] Data information acquisition module, which is used to acquire data center information and transmit the acquired data center information to the model optimization and establishment module;

[0029] The model optimization and establishment module is used to determine the tuning target based on the acquired data center information, select the appropriate model based on the tuning target, train and optimize the model to obtain the AI ​​large model, and transmit the generated AI large model to the server status analysis module;

[0030] The server status analysis module is used to substitute the real-time data center energy consumption into the AI ​​large model for prediction to obtain predicted data, and compare the predicted data with the normal value to generate a comparison result, which includes normal results and abnormal results. At the same time, the abnormal results are analyzed, and the server classification results are obtained by classifying different servers according to their real-time energy consumption. For high-energy-consuming servers, the current real-time processing tasks are scheduled and optimized by combining historical data, generating tasks to be processed and transmitting them to the tuning analysis processing module;

[0031] A tuning analysis processing module is used to obtain pending tasks and low-energy servers, determine the real-time load of the low-energy servers, perform scheduling optimization processing on the low-energy servers that meet the real-time load requirements, generate tuning control information, adjust the low-energy servers according to the tuning control information, and perform tuning analysis on the CPU frequency and voltage parameters of the servers according to the actual workload of the adjusted servers, generate secondary tuning information, and transmit the generated secondary tuning information and tuning control information to the processing information output module;

[0032] The processing information output module is used to display the acquired secondary tuning information and tuning control information to the corresponding operator.

[0033] This invention provides a data center energy-saving optimization control method and system based on an AI large model. Compared with the existing technology, it has the following advantages:

[0034] The present invention uses a large AI model to accurately predict the energy consumption of the data center, can timely adjust the equipment operating parameters according to the prediction results, realize refined energy consumption management, can interact with the data center environment in real time, can adjust the equipment operating parameters in real time, dynamically adapt to changes in business load, and improve energy utilization efficiency. At the same time, through the reasonable scheduling and optimization of tasks and servers, the tasks to be processed are reasonably allocated to low-energy consumption servers, further improving the overall energy-saving effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Figure 1 It is a diagram of the steps and methods of the present invention;

[0036] Figure 2 This is a block diagram of the system principle of the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] For example 1, please refer to Figure 1 , this application provides a data center energy-saving tuning control method based on an AI large model, the method specifically comprising the following steps:

[0039] Step S001: Acquire data center information and determine the tuning target. Select a suitable model based on the tuning target, and train and optimize the model to obtain an AI large model.

[0040] Obtain data center information, then determine the data center characteristics and energy-saving targets based on the data center information, and simultaneously obtain a learning model suitable for the data center characteristics and a learning model corresponding to the energy-saving targets, and select a common model that is suitable for both, and use this as the standard to determine the selected model.

[0041] Specifically, if the data center mostly generates structured data such as equipment operating parameters and energy consumption data, traditional machine learning models such as linear regression, decision trees, and random forests can be selected; for complex nonlinear relationships, deep neural networks (DNNs) can be selected. When it contains a large amount of unstructured data such as video surveillance data and text logs, convolutional neural networks (CNNs) are suitable for processing image and video data, and recurrent neural networks (RNNs) and their variants such as LSTM and GRU perform well in processing sequential text data. If the goal is to adjust equipment operating parameters in real time to achieve energy saving, reinforcement learning models such as deep Q networks (DQNs) and their extensions can continuously learn optimal strategies through interaction with the environment and make decisions in real time. For example, if a data center mainly uses structured data and hopes to adjust equipment operating parameters in real time to achieve energy saving, then it can be considered to combine the random forest model with the reinforcement learning model, use the random forest model to analyze the operating status of the data center, provide a decision basis for the reinforcement learning model, and thus more effectively achieve energy saving goals.

[0042] Then, the model is trained on the preprocessed data based on the selected model, and the AI ​​large model is obtained by adjusting the model parameters and hyperparameters, such as learning rate, number of layers, number of neurons, etc., to optimize the performance of the model so that it can accurately predict the energy consumption and operating status of the data center.

[0043] In step S002, the real-time data center energy consumption is substituted into the AI ​​big model for prediction to obtain predicted data, and the predicted data is compared with the normal value to generate a comparison result, and the comparison result includes a normal result and an abnormal result.

[0044] The real-time energy consumption corresponding to the data center is obtained, and the real-time energy consumption here is expressed as the energy consumption data corresponding to different servers, including CPU utilization, memory utilization, and voltage parameters, etc. The obtained real-time energy consumption is substituted into the AI ​​large model for analysis to obtain predicted data. At the same time, the obtained predicted data is compared and analyzed with the normal value, and the specific value of the normal value is set by the operator. If the predicted data is greater than the normal value, it means that the overall predicted energy consumption of the corresponding data center exceeds the value under normal circumstances, and an abnormal result is generated. On the contrary, if the predicted data is less than the normal value, it means that the predicted energy consumption corresponding to the entire data center does not exceed the value under normal circumstances, and a normal result is generated.

[0045] For example, for Server A, the normal range for CPU utilization is set at 40% to 70%, the normal range for memory utilization is set at 30% to 65%, and the normal range for voltage parameters is set at 11V to 13V. The predicted data output by the AI ​​large model is carefully compared and analyzed with these set normal values. For example, if Server A's predicted CPU utilization of 80% exceeds the upper limit of the normal range of 70%, this indicates that the corresponding data center's overall predicted energy consumption exceeds the normal value, and the system will immediately generate abnormal result feedback.

[0046] Step S003: Analyze the abnormal results obtained, classify the servers according to their real-time energy consumption to obtain server classification results, and optimize the current real-time processing tasks by combining historical data with high-energy-consuming servers to generate tasks to be processed.

[0047] The server corresponding to the abnormal result is recorded as an abnormal server. At the same time, the real-time energy consumption of the abnormal server is obtained, and the real-time energy consumption is compared with the preset energy consumption value. The value of the preset energy consumption value is set by the operator. If the real-time energy consumption is greater than the preset energy consumption value, the abnormal server is classified as a high-energy consumption server. Conversely, if the real-time energy consumption is less than the preset energy consumption value, the abnormal server is classified as a low-energy consumption server. Specifically, the preset energy consumption value here is greater than the normal value, and the low-energy consumption server is a secondary classification case among the abnormal servers.

[0048] Then, all high-energy-consuming servers are obtained, labeled as i, and i = 1, 2, ..., j, where j represents the number of high-energy-consuming servers. At the same time, the real-time processing tasks of the high-energy-consuming servers are obtained, and the high-load tasks and low-load tasks are classified according to the energy consumption requirements of the real-time processing tasks. Here, the same analysis and classification processing is performed on all high-energy-consuming servers, wherein the energy consumption requirements are compared with the corresponding comparison values, and the comparison values ​​are determined based on the CPU utilization of the server itself. Then, the historical data of the high-energy-consuming servers is obtained, and all processed high-load task records in the historical data are obtained, and the corresponding processing times are obtained. The processing times obtained here are the processing times corresponding to different high-load tasks in different historical data. At the same time, the processing proportion values ​​corresponding to different high-load tasks are calculated, and the processing proportion values ​​are sorted from large to small.

[0049] The sorted processing ratio value is matched with the high-load task. If the match is successful, the corresponding high-load task is marked as a task to be processed. If the match is unsuccessful, it is not processed. The successful match here is represented by the corresponding number of times the corresponding high-load task exists in the sorting information.

[0050] Assume that the data center has identified three high-energy-consuming servers through the above steps, numbered 1, 2, and 3 respectively. For the high-energy-consuming server numbered 2, historical data shows that it has processed high-load task A 10 times and task B 5 times. If server 2 has processed a total of 15 high-load tasks (task A 10 times, task B 5 times), then the processing ratio of task A is 10÷15≈66.7%, and the processing ratio of task B is 5÷15≈33.3%. After sorting, the processing ratio of task A is 66.7%, which corresponds to the sorting information, so task A is marked as a pending task; and if there is a new high-load task C, its processing ratio does not exist in the sorting information, then task C will not be processed.

[0051] Step S004: Obtain tasks to be processed and low-energy consumption servers, determine the real-time load of the low-energy consumption servers, and perform scheduling optimization processing on the low-energy consumption servers that meet the real-time load requirements to generate tuning control information.

[0052] Obtain all low-energy servers and label them as n, and n = 1, 2, ..., m, where m is the number of low-energy servers, and obtain the real-time load of the low-energy servers, and at the same time obtain the label of the pending tasks as a, and a = 1, 2, ..., b, where b represents the number of pending tasks, then sort the pending tasks a from large to small according to energy consumption requirements, and calculate the sum of the energy consumption requirements of the pending tasks and the real-time load of the low-energy servers, and at the same time compare the calculated sum with the judgment load value, and the judgment load value is set according to the load situation of the low-energy server itself, if the sum is greater than the judgment load value, then reallocate the pending tasks, if the sum is less than the judgment value, then directly perform information scheduling optimization processing, and generate tuning control information.

[0053] For example, the energy consumption requirement of pending task 1 is 30% of CPU resources, the energy consumption requirement of task 2 is 25% of CPU resources, the energy consumption requirement of task 3 is 35% of CPU resources, the energy consumption requirement of task 4 is 20% of CPU resources, and the energy consumption requirement of task 5 is 15% of CPU resources. After sorting, the order of tasks is 3, 1, 2, 4, 5. For example, the real-time CPU load of low-energy server 1 is 40%, and the memory load is 30%. For example, for low-energy server 1, based on its performance and historical data, the CPU load corresponding to the load value is set to 70%. Taking low-energy server 1 and pending task 3 as an example, the energy consumption requirement of task 3 is 35% of CPU resources, and the real-time CPU load of low-energy server 1 is 40%. Then their sum is 40% + 35% = 75%.

[0054] The sum of the low-energy server 1 and the pending task 3, 75%, is greater than the judgment load value of 70%. This indicates that if the task is assigned to this low-energy server, it may cause the server load to be too high, affecting its normal operation and even causing new energy consumption anomalies or performance problems. For example, the energy consumption requirement of the pending task 5 is 15% of the CPU resources, which is added to the real-time CPU load of 40% of the low-energy server 1. The sum is 40% + 15% = 55%, which is less than the judgment load value of 70%. In this case, the task can be directly subjected to information scheduling optimization processing.

[0055] Step S005 , adjusting the low-energy consumption server according to the tuning control information, and performing tuning analysis on the CPU frequency and voltage parameters of the server according to the actual workload of the adjusted server to generate secondary tuning information.

[0056] The low-energy-consuming server is adjusted using the tuning control information, and the adjusted server is recorded as the optimized server. The actual workload corresponding to the optimized server is then obtained, and the workload is graded. The collected workload data is graded according to pre-set standards or rules determined by a machine learning algorithm. For example, the workload is divided into three levels: low, medium, and high. The classification standards are: CPU utilization rate less than 40% and memory usage less than 30GB is a low load level; CPU utilization rate between 40% and 70% and memory usage between 30GB and 60GB is a medium load level; CPU utilization rate greater than 70% and memory usage greater than 60GB is a high load level. According to this standard, the workload of optimized server 1 belongs to the medium load level, and the CPU frequency and voltage parameters are determined according to different levels to generate a parameter adjustment range. The actual workload is then matched with the parameter adjustment range to obtain the corresponding CPU frequency and voltage parameters, and the satisfaction of the task requirements corresponding to the optimized server is determined.

[0057] For example, for low-load workloads, in order to achieve energy saving, the CPU frequency adjustment range is set to 1.0GHz-1.5GHz, and the voltage adjustment range is set to 0.8V-1.0V; for medium-load levels, such as the case of optimized server 1, the CPU frequency adjustment range is set to 1.5GHz-2.0GHz, and the voltage adjustment range is set to 1.0V-1.2V; for high-load levels, in order to ensure server performance, the CPU frequency adjustment range is set to 2.0GHz-2.5GHz, and the voltage adjustment range is set to 1.2V-1.4V. The current CPU utilization rate of optimized server 1 is 50%, the memory usage is 40GB, and the I / O read and write frequency is 100 times per second. Since its workload is medium-load level, the corresponding CPU frequency parameter range is 1.5GHz-2.0GHz, and the voltage parameter range is 1.0V-1.2V.

[0058] If the adjusted optimized server can meet the corresponding task requirements, the optimized server will be adjusted based on the corresponding CPU frequency and voltage parameters. Otherwise, if the task requirements are not met, the corresponding parameter adjustment range will be increased, and secondary optimization information will be generated at the same time.

[0059] For example, the task currently undertaken by the optimization server 1 requires a certain amount of data processing work to be completed within 1 hour. Through analysis and evaluation of the server's current performance indicators and task volume, if the server is configured according to the current CPU frequency and voltage parameters, it is expected that the server can complete the task within the specified time, that is, it is determined to be able to meet the task requirements.

[0060] If the adjusted optimized server can meet the corresponding task requirements, for example, if optimized server 1 meets the task requirements, then the parameters of the optimized server are adjusted based on the matched CPU frequency and voltage parameters. In other words, the configuration of optimized server 1 is adjusted to the CPU frequency range of 1.5GHz-2.0GHz and the voltage range of 1.0V-1.2V to achieve a balance between energy saving and performance.

[0061] Conversely, if the evaluation finds that the optimized server cannot meet the task requirements, for example, if optimized server 1 is not expected to complete the task within 1 hour, the corresponding parameter adjustment range needs to be increased. For example, the CPU frequency adjustment range can be increased to 2.0GHz-2.5GHz, and the voltage adjustment range can be increased to 1.2V-1.4V. At the same time, secondary optimization information is generated.

[0062] For example 2, please refer to Figure 2 This application provides a data center energy-saving tuning control system based on AI big model, including data information acquisition module, model optimization establishment module, server status analysis module, tuning analysis processing module and processing information output module, combined with Figure 2 It can be known that the functional modules are electrically connected in a unidirectional manner.

[0063] Data information acquisition module, which is used to acquire data center information and transmit the acquired data center information to the model optimization and establishment module;

[0064] A model optimization and establishment module is used to determine the tuning target based on the acquired data center information, select a suitable model based on the tuning target, train and optimize the model to obtain an AI large model, and transmit the generated AI large model to the server status analysis module. The specific processing method is similar to the processing process of step S001 in Example 1;

[0065] A server status analysis module is used to substitute the real-time data center energy consumption into the AI ​​large model for prediction to obtain predicted data, and compare the predicted data with the normal value to generate a comparison result, and the comparison result includes a normal result and an abnormal result. The specific processing method is the same as the processing process of step S002 in Example 1. At the same time, the abnormal results obtained are analyzed, and server classification results are obtained according to the real-time energy consumption of different servers. For high-energy consumption servers, the current real-time processing tasks are scheduled and optimized by combining historical data to generate tasks to be processed, and the tasks are transmitted to the tuning analysis processing module. The specific processing method is the same as the processing process of step S003 in Example 1;

[0066] A tuning analysis processing module is used to obtain tasks to be processed and low-energy servers, judge the real-time load of the low-energy servers, and perform scheduling optimization processing on the low-energy servers that meet the real-time load requirements to generate tuning control information. The specific processing method is similar to the processing process of step S004 in Example 1. The low-energy servers are adjusted according to the tuning control information, and the CPU frequency and voltage parameters of the servers are tuned and analyzed according to the actual workload of the adjusted servers to generate secondary tuning information. The generated secondary tuning information and tuning control information are transmitted to the processing information output module. The specific processing method is similar to the processing process of step S005 in Example 1.

[0067] The processing information output module is used to display the acquired secondary tuning information and tuning control information to the corresponding operator.

[0068] Some of the data in the above formulas are numerically calculated based on their dimension values. Meanwhile, the contents not described in detail in this specification belong to the prior art known to those skilled in the art.

[0069] The above embodiments are only used to illustrate the technical method of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A data center energy-saving tuning control method based on AI big model, characterized in that: The method specifically comprises the following steps: Step S001, acquiring data center information and determining tuning targets, selecting a suitable model according to the tuning targets, and training and optimizing the model to obtain an AI large model; Step S002: Substitute the real-time data center energy consumption into the AI ​​big model for prediction to obtain prediction data, and compare the prediction data with the normal value to generate a comparison result, and the comparison result includes a normal result and an abnormal result; Step S003, analyzing the abnormal results obtained, classifying different servers according to their real-time energy consumption to obtain server classification results, and optimizing the current real-time processing tasks by combining historical data with high-energy-consuming servers to generate tasks to be processed; Step S004, obtaining tasks to be processed and low-energy consumption servers, determining the real-time load of the low-energy consumption servers, and performing scheduling optimization processing on the low-energy consumption servers that meet the real-time load requirements to generate tuning control information; Step S005, adjusting the low energy consumption server according to the tuning control information, and performing tuning analysis on the CPU frequency and voltage parameters of the server according to the actual workload of the adjusted server to generate secondary tuning information.

2. According to the data center energy-saving tuning control method based on AI big model according to claim 1, it is characterized in that: The specific method of obtaining the AI ​​large model in step S001 is: Obtain data center information, then determine the characteristics of the data center and energy-saving targets based on the data center information, and simultaneously obtain a learning model suitable for the characteristics of the data center and a learning model corresponding to the energy-saving targets, and select a common model that is suitable for both. Use this as the standard to determine the selected model, and then train the model on the preprocessed data based on the determined selected model. Optimize the model by adjusting the model's parameters and hyperparameters to obtain a large AI model.

3. According to the data center energy-saving tuning control method based on AI big model in claim 1, it is characterized in that: The specific method of generating the comparison result in step S002 is: The real-time energy consumption corresponding to the data center is obtained, and the obtained real-time energy consumption is substituted into the AI ​​large model for analysis to obtain predicted data. At the same time, the obtained predicted data is compared and analyzed with the normal value. If the predicted data is greater than the normal value, an abnormal result is generated. Conversely, if the predicted data is less than the normal value, a normal result is generated.

4. According to the data center energy-saving tuning control method based on AI big model according to claim 1, it is characterized in that: The specific method of analyzing the abnormal results in step S003 is: The server corresponding to the abnormal result is recorded as an abnormal server. At the same time, the real-time energy consumption of the abnormal server is obtained and compared with the preset energy consumption value. If the real-time energy consumption is greater than the preset energy consumption value, it is classified as a high-energy consumption server, otherwise it is classified as a low-energy consumption server; Then, all high-energy-consuming servers are labeled as i, and i=1, 2, ..., j, where j represents the number of high-energy-consuming servers. At the same time, the real-time processing tasks of the high-energy-consuming servers are obtained, and classified into high-load tasks and low-load tasks according to the energy consumption requirements of the real-time processing tasks, and then the two are analyzed.

5. According to the data center energy-saving tuning control method based on AI big model according to claim 4, it is characterized in that: The specific method of analyzing the high-load task and the low-load task in step S003 is: Obtain the historical data of high-energy consumption servers, obtain all processed high-load task records in the historical data, and obtain the corresponding processing times, and calculate the processing proportion values ​​corresponding to different high-load tasks, and sort the processing proportion values ​​from large to small; The sorted processing ratio values ​​are matched with the high-load tasks. If the match is successful, the corresponding high-load tasks are marked as pending tasks. If the match is unsuccessful, they are not processed.

6. According to the data center energy-saving optimization control method based on AI big model in claim 1, it is characterized in that: The specific method of generating the tuning control information in step S004 is: Obtain all low-energy servers and label them as n, where n=1, 2, ..., m, where m is the number of low-energy servers, and obtain the real-time load of the low-energy servers. At the same time, obtain the pending tasks and label them as a, where a=1, 2, ..., b, where b represents the number of pending tasks. Then sort the pending tasks a from large to small according to the energy demand, and calculate the sum of the energy demand of the pending tasks and the real-time load of the low-energy servers. At the same time, compare the calculated sum with the judgment load value. If the sum is greater than the judgment load value, reallocate the pending tasks. If the sum is less than the judgment value, directly perform information scheduling optimization processing and generate tuning control information.

7. According to claim 1, a data center energy-saving tuning control method based on AI big model is characterized in that: The specific method of generating the secondary tuning information in step S005 is: The low-energy consumption server is adjusted with the tuning control information, and the adjusted server is recorded as the optimized server. Then, the actual workload corresponding to the optimized server is obtained, and the workload is graded. The CPU frequency and voltage parameters are determined according to different grades, and the parameter adjustment interval is generated. Then, the actual workload is matched with the parameter adjustment interval to obtain the corresponding CPU frequency and voltage parameters, and the task requirements corresponding to the optimized server are judged. If the adjusted optimized server can meet the corresponding task requirements, the optimized server will be adjusted based on the corresponding CPU frequency and voltage parameters. Otherwise, if it does not meet the task requirements, the corresponding parameter adjustment range will be increased, and secondary optimization information will be generated at the same time.

8. A data center energy-saving optimization control system based on an AI big model, used to execute a data center energy-saving optimization control method based on an AI big model according to any one of claims 1 to 7, characterized in that: It includes data information acquisition module, model optimization establishment module, server status analysis module, tuning analysis processing module and processing information output module; A data information acquisition module, which is used to acquire data center information and transmit the acquired data center information to the model optimization establishment module; Model optimization establishment module, which is used to determine the tuning target based on the acquired data center information, select the appropriate model based on the tuning target, train and optimize the model to obtain the AI ​​big model, and transmit the generated AI big model to the server status analysis module; Server status analysis module: This module is used to substitute the real-time data center energy consumption into the AI ​​big model for prediction to obtain prediction data, and compare the prediction data with the normal value to generate a comparison result, and the comparison result includes a normal result and an abnormal result. At the same time, the abnormal result is analyzed, and the server classification result is obtained by classifying the real-time energy consumption of different servers. For high-energy consumption servers, the current real-time processing tasks are scheduled and optimized by combining historical data, generating tasks to be processed, and transmitting them to the tuning analysis processing module; The tuning analysis processing module is used to obtain pending tasks and low-energy servers, judge the real-time load of the low-energy servers, perform scheduling optimization processing on the low-energy servers that meet the real-time load requirements, generate tuning control information, adjust the low-energy servers according to the tuning control information, and perform tuning analysis on the CPU frequency and voltage parameters of the servers according to the actual workload of the adjusted servers, generate secondary tuning information, and transmit the generated secondary tuning information and tuning control information to the processing information output module; The processing information output module is used to display the acquired secondary tuning information and tuning control information to the corresponding operator.

Citation Information

Patent Citations

  • Energy-saving control method and device for data center, electronic equipment and readable medium

    CN116361703A

  • Data center energy-saving control method based on AI algorithm

    CN117850494A

Cited By

  • Energy efficiency evaluation method of artificial intelligence data center

    CN120973622A