Distributed computing power center load and energy consumption cooperative scheduling method based on artificial intelligence

By building an artificial intelligence prediction model and generating scheduling strategies in the distributed computing power center, the problems of unbalanced workload and excessive energy consumption are solved, and efficient resource utilization and energy consumption reduction are achieved.

CN120216197APending Publication Date: 2025-06-27HEFEI UNIV OF TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510367674.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The problems of unbalanced workloads and excessive energy consumption of distributed computing power centers lead to unreasonable resource allocation, making it difficult to achieve efficient resource utilization and significant reduction in energy consumption.

Method used

Using an artificial intelligence-based approach, workload prediction models, energy consumption prediction models and cooling system energy consumption prediction models are constructed through GAT and LSTM, and work together to establish a quantitative relationship between workload and energy consumption, and use generative artificial intelligence to generate flexible and efficient scheduling strategies.

Benefits of technology

It significantly improves the flexibility and intelligence of the system, realizes effective coordinated scheduling of workloads and energy consumption, improves resource utilization, reduces energy consumption, and adapts to complex and changeable computing power demand scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216197A_ABST
    Figure CN120216197A_ABST
Patent Text Reader

Abstract

The invention provides a distributed computing power center load and energy consumption collaborative scheduling method and system based on artificial intelligence, a storage medium and electronic equipment, and relates to the technical field of load and energy consumption collaborative scheduling. According to the method, a computing power center node workload prediction model, a computing power center node workload energy consumption prediction model and a cooling system energy consumption prediction model are respectively constructed by utilizing GAT and LSTM, the three models cooperatively work through a cascade architecture, and the output of a preorder model is used as the input of a subsequent model, so that the data interaction capability among the models is enhanced; and the overall prediction performance is improved. Besides, the strong language generation and knowledge reasoning capabilities of the generative artificial intelligence are utilized to generate a flexible and efficient scheduling strategy, so that the flexibility and the intelligence degree of the system are remarkably improved, and the system can better adapt to complex and changeable computing power demand scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of collaborative scheduling of load and energy consumption, and particularly to a method, system, storage medium and electronic device for collaborative scheduling of load and energy consumption in a distributed computing power center based on artificial intelligence. Background Art

[0002] In the wave of digital transformation, distributed computing power centers have become key infrastructures for promoting the development of various industries. Distributed computing power centers face severe challenges of unbalanced workload and excessive energy consumption during operation.

[0003] On the one hand, due to the huge differences in computing power requirements of different applications, and the random arrival time and duration of tasks, the workload distribution among computing nodes is extremely uneven. Some nodes may be overloaded due to taking on too many tasks, resulting in performance degradation or even collapse. While other nodes may be idle, causing waste of resources. On the other hand, the energy consumption problem of the computing power center cannot be ignored. Computing nodes consume a large amount of electrical energy during operation, and at the same time, in order to maintain their normal working temperature, the cooling system also needs to consume a large amount of energy. Excessive energy consumption not only increases the operating cost, but also causes great pressure on the environment. Therefore, realizing the collaborative scheduling of workload and energy consumption in distributed computing power centers, improving resource utilization rate, and reducing energy consumption have become key issues that need to be solved urgently.

[0004] Taking the edge computing power center as an example, as a typical application scenario of distributed computing power centers, its nodes are widely distributed and resource-constrained. The dynamic nature of task workloads and the complexity of the network environment further exacerbate the problems of unbalanced workload and excessive energy consumption. For example, in smart cities or industrial Internet, edge nodes need to process a large amount of local data in real time. However, due to uneven task allocation, some nodes may not be able to meet real-time requirements due to overload, while other nodes are operating inefficiently. At the same time, the energy supply of edge devices is usually limited, and excessive energy consumption will significantly shorten the device life and increase the operation and maintenance cost.

[0005] Currently, most methods only focus on the research of either the workload or energy consumption of the computing power center alone, while ignoring the strong correlation among the workload of the computing power center, the energy consumption of the computing power center workload, and the energy consumption of the cooling system. As a result, it is difficult to accurately identify the hot spots and flow paths of energy consumption during the energy consumption management and optimization process, and thus it is impossible to make targeted improvements. Summary of the Invention

[0006] (I) Technical Problems to be Solved

[0007] In view of the deficiencies of the prior art, the present invention provides a method, system, storage medium and electronic device for collaborative scheduling of load and energy consumption of a distributed computing power center based on artificial intelligence, which solves the technical problem of unreasonable resource allocation caused by separately scheduling the workload or energy consumption of the computing power center.

[0008] (2) Technical solution

[0009] To achieve the above objectives, the present invention is realized through the following technical solutions:

[0010] A method for collaborative scheduling of load and energy consumption of a distributed computing power center based on artificial intelligence, including:

[0011] Collect and preprocess the historical data of node workloads, the historical data of workload requirements of new computing power tasks, the historical energy consumption data of node workloads, the historical data of node environments, and the historical energy consumption data of the cooling system;

[0012] Based on the preprocessed historical data of node workloads and the historical data of workload requirements of new computing power tasks, divide the first training set and the first test set, and use GAT and LSTM to construct a computing power center node workload prediction model to obtain node workload prediction data;

[0013] Based on the preprocessed historical energy consumption data of node workloads and the historical data of node environments, combined with the node workload prediction data, divide the second training set and the second test set, and use GAT and LSTM to construct a computing power center node workload energy consumption prediction model to obtain the energy consumption prediction data of node workloads;

[0014] Based on the preprocessed historical energy consumption data of the cooling system and the historical data of node environments, combined with the energy consumption prediction data of node workloads, divide the third training set and the third test set, and use LSTM to construct an energy consumption prediction model for the cooling system;

[0015] Collect the workload requirement data of new computing power tasks, as well as the latest historical data of node workloads, workload energy consumption, node environments, and the latest historical energy consumption of the cooling system, and predict the workloads, energy consumption of workloads, and energy consumption data of the cooling system of nodes within a future time window through the computing power center node workload prediction model, the computing power center node workload energy consumption prediction model, and the energy consumption prediction model of the cooling system;

[0016] Fill the workloads, energy consumption of workloads, and energy consumption data of the cooling system within the future time window into a preset prompt information template and use it as the input of generative artificial intelligence to obtain the final collaborative scheduling plan for the workload - energy consumption of the distributed computing power center.

[0017] Preferably, the computing power center node workload prediction model includes two LSTM layers and two GAT layers; where:

[0018] The first LSTM layer is used to receive the node workload historical data of multiple historical time steps of multiple computing nodes, and transfer the features output by it to the first GAT layer to obtain the first feature through multiple attention heads;

[0019] The second LSTM layer is used to receive the historical data of the workload requirements of multiple new computing tasks, and after reshaping the features output by it, obtain the second feature;

[0020] The second GAT layer is used to receive the combined result of the first feature and the second feature, and obtain the node workload prediction data of multiple historical time steps of multiple computing nodes.

[0021] Preferably, the computing power center node workload energy consumption prediction model includes three LSTM layers and two GAT layers; where:

[0022] The third LSTM layer is used to receive the node workload prediction data of multiple historical time steps of multiple computing nodes to obtain the third feature;

[0023] The fourth LSTM layer is used to receive the energy consumption historical data of the node workload of multiple historical time steps of multiple computing nodes to obtain the fourth feature;

[0024] The fifth LSTM layer is used to receive the node environment historical data of multiple historical time steps, and after reshaping the features output by it, obtain the fifth feature;

[0025] The third GAT layer is used to receive the combined result of the third feature, the fourth feature and the fifth feature, and transfer it to the fourth GAT layer to obtain the energy consumption prediction data of the node workload of multiple historical time steps of multiple computing nodes.

[0026] Preferably, the LSTM-based cooling system energy consumption prediction model includes two LSTMs; where:

[0027] After reshaping the energy consumption prediction data of the node workload of multiple historical time steps of multiple computing nodes, obtain the sixth feature; and use the energy consumption historical data of the cooling system as the seventh feature, and the preprocessed node environment historical data as the eighth feature;

[0028] The sixth LSTM layer is used to receive the combined result of the sixth feature, the seventh feature and the eighth feature, and transfer it to the seventh LSTM layer to obtain the energy consumption prediction data of multiple historical time steps of the cooling system.

[0029] Preferably, the generative artificial intelligence uses GPT-4.

[0030] Preferably, it is characterized in that

[0031] The node workload historical data includes any one or a combination of any several of the CPU usage rate, average load, memory occupancy, memory free amount, memory usage rate, disk read and write speed, disk read and write times, disk I / O waiting time, number of running processes, task queue length, CPU time occupied by each process, number of transactions processed per second, and data volume transmitted per second.

[0032] Preferably, the historical data of the workload requirements of the new computing power tasks includes any one or a combination of any several of the task type, task priority, data volume size, model complexity, number of users, frequency of user operations, number of instructions to be executed per second, number of floating-point operations per second, frequency of data reading and writing, and task parallelism;

[0033] Preferably, the historical energy consumption data of the node workload includes any one or a combination of any several of the overall energy consumption of the node within the input step, the average energy consumption of the node within the input step, the energy consumption of the CPU, the energy consumption of the memory, and the energy consumption of the storage device.

[0034] Preferably, the node environment historical data includes any one or a combination of any several of the node environment temperature, the computer room temperature of the computing power center node, the difference between the weather temperature and the node computer room temperature, the environmental humidity, and the computer room humidity.

[0035] Preferably, the historical energy consumption data of the cooling system includes any one or a combination of any several of the overall energy consumption of the node cooling system within the input step and the average energy consumption of the node cooling system within the input step.

[0036] A distributed computing power center load and energy consumption collaborative scheduling system based on artificial intelligence, comprising:

[0037] A data collection and preprocessing module for collecting and preprocessing node workload historical data, historical data of workload requirements of new computing power tasks, historical energy consumption data of node workloads, node environment historical data, and historical energy consumption data of the cooling system;

[0038] A model training module for dividing a first training set and a first test set based on the preprocessed node workload historical data and the historical data of the workload requirements of the new computing power tasks, and constructing a computing power center node workload prediction model using GAT and LSTM to obtain node workload prediction data;

[0039] Based on the historical energy consumption data of the node workload and the historical node environment data after preprocessing, combined with the predicted data of the node workload, divide the second training set and the second test set, and use GAT and LSTM to construct a computing power center node workload energy consumption prediction model to obtain the predicted energy consumption data of the node workload;

[0040] And based on the historical energy consumption data of the cooling system and the historical node environment data after preprocessing, combined with the predicted energy consumption data of the node workload, divide the third training set and the third test set, and use LSTM to construct an energy consumption prediction model for the cooling system;

[0041] The data prediction module is used to collect the workload demand data of new computing power tasks, as well as the latest historical data of node workload, workload energy consumption, node environment, and the latest historical data of energy consumption of the cooling system. Through the computing power center node workload prediction model, the computing power center node workload energy consumption prediction model, and the energy consumption prediction model of the cooling system, predict the workload, energy consumption of the workload, and energy consumption data of the cooling system of the node within the future time window;

[0042] The solution generation module is used to fill the workload, energy consumption of the workload, and energy consumption data of the cooling system within the future time window into a preset prompt information template, and use it as the input of the generative artificial intelligence to obtain the final distributed computing power center workload - energy consumption collaborative scheduling solution.

[0043] A storage medium stores a computer program for the collaborative scheduling of the load and energy consumption of a distributed computing power center based on artificial intelligence, wherein the computer program enables a computer to execute the above - mentioned method for the collaborative scheduling of the load and energy consumption of the distributed computing power center.

[0044] An electronic device includes:

[0045] One or more processors; a memory; and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include those for executing the above - mentioned method for the collaborative scheduling of the load and energy consumption of the distributed computing power center.

[0046] (III) Advantageous Effects

[0047] The present invention provides a method, system, storage medium, and electronic device for the collaborative scheduling of the load and energy consumption of a distributed computing power center based on artificial intelligence. Compared with the prior art, the following advantageous effects are achieved:

[0048] In the present invention, the Graph Attention Network (GAT) and Long Short-Term Memory (LSTM) are respectively used to construct a workload prediction model for computing power center nodes, a workload energy consumption prediction model for computing power center nodes, and an energy consumption prediction model for the cooling system. The three work together to establish a quantitative relationship between workload and energy consumption, realizing data sharing and interaction. In addition, by leveraging the powerful language generation and knowledge reasoning capabilities of generative artificial intelligence, a flexible and efficient scheduling strategy is generated, significantly improving the flexibility and intelligence of the system and better adapting to complex and changing computing power demand scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0050] Figure 1 FIG. [X] is a block diagram of a distributed computing power center load and energy consumption collaborative scheduling method based on artificial intelligence provided by an embodiment of the present invention;

[0051] Figure 2 FIG. [X] is a flowchart of a distributed computing power center load and energy consumption collaborative scheduling method based on artificial intelligence according to an embodiment of the present invention;

[0052] Figure 3 FIG. [X] is a flowchart for constructing a workload prediction model for computing power center nodes according to an embodiment of the present invention;

[0053] Figure 4 FIG. [X] is a flowchart for constructing a workload energy consumption prediction model for computing power center nodes according to an embodiment of the present invention;

[0054] Figure 5 FIG. [X] is a flowchart for constructing an energy consumption prediction model for the cooling system according to an embodiment of the present invention;

[0055] Figure 6 FIG. [X] is a block diagram of a distributed computing power center load and energy consumption collaborative scheduling system based on artificial intelligence provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0056] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are clearly and completely described below. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0057] Embodiments of this application provide a method, system, storage medium, and electronic device for collaborative scheduling of load and energy consumption in a distributed computing power center based on artificial intelligence, solving the technical problem of unreasonable resource allocation caused by separately scheduling the workload or energy consumption of the computing power center.

[0058] The key points of the technical solution in the embodiments of this application for solving the above technical problems are as follows:

[0059] 1. Graph structure modeling: GAT can model the nodes and their mutual relationships in a distributed computing power center as a graph structure, and capture the complex correlation relationships between nodes through the graph attention mechanism.

[0060] 2. Node feature fusion: In the energy consumption prediction module, GAT can effectively fuse features from multiple aspects such as node workload and node environment, and can also comprehensively consider the workload of adjacent nodes and the communication link status between nodes, so as to more accurately predict the changing trends of node workload and energy consumption.

[0061] 3. Workload-energy consumption correlation analysis: The node workload prediction module, the energy consumption prediction module of the node workload, and the energy consumption prediction module of the cooling system cooperate with each other. By deeply analyzing the workload and energy consumption data, a quantitative relationship between workload and energy consumption is established.

[0062] 4. Generation of intelligent scheduling scheme: Based on the powerful language generation ability and knowledge reasoning ability of generative artificial intelligence, and according to the prompt information constructed from the workload and energy consumption prediction results, an optimal workload-energy consumption collaborative scheduling scheme is generated.

[0063] 5. Adaptive adjustment: Generative artificial intelligence can dynamically adjust the scheduling scheme according to the real-time workload and energy consumption of the computing power center to adapt to changing requirements, ensuring that the computing power center is always in an efficient and energy-saving operating state.

[0064] In addition, explanations of the terms involved in the embodiments of the present invention are supplemented:

[0065] 1) Graph Attention Network (abbreviated as GAT) is a neural network model based on graph structure, aiming to solve some limitations of Graph Convolutional Network (GCN).

[0066] 2) Long short-term memory (abbreviated as LSTM) is a special Recurrent Neural Network (RNN), mainly to solve the problems of gradient vanishing and gradient explosion in the training process of long sequences.

[0067] To better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings of the specification and specific embodiments.

[0068] Example 1:

[0069] As Figure 1 shown, the embodiment of the present invention provides a method for collaborative scheduling of load and energy consumption of a distributed computing power center based on artificial intelligence, including:

[0070] S1. Collect and preprocess the historical data of node workload, the historical data of workload requirements of new computing power tasks, the historical energy consumption data of node workload, the historical data of node environment, and the historical energy consumption data of the cooling system;

[0071] S2. Based on the preprocessed historical data of node workload and the historical data of workload requirements of new computing power tasks, divide the first training set and the first test set, and use GAT and LSTM to construct a computing power center node workload prediction model to obtain node workload prediction data;

[0072] S3. Based on the preprocessed historical energy consumption data of node workload and the historical data of node environment, combine the node workload prediction data, divide the second training set and the second test set, and use GAT and LSTM to construct a computing power center node workload energy consumption prediction model to obtain the energy consumption prediction data of node workload;

[0073] S4. Based on the preprocessed historical energy consumption data of the cooling system and the historical data of node environment, combine the energy consumption prediction data of node workload, divide the third training set and the third test set, and use LSTM to construct an energy consumption prediction model of the cooling system;

[0074] S5. Collect the workload requirement data of new computing power tasks, as well as the latest historical data of node workload, workload energy consumption, node environment, and the latest historical energy consumption of the cooling system, and predict the workload, workload energy consumption, and energy consumption data of the cooling system of the node within the future time window through the computing power center node workload prediction model, the computing power center node workload energy consumption prediction model, and the energy consumption prediction model of the cooling system;

[0075] S6. Fill the workload, workload energy consumption, and energy consumption data of the cooling system within the future time window into a preset prompt information template, and use it as the input of generative artificial intelligence to obtain the final collaborative scheduling scheme for the workload - energy consumption of the distributed computing power center.

[0076] In the embodiments of the present invention, the Graph Attention Network (GAT) and Long Short-Term Memory (LSTM) are respectively used to construct a workload prediction model for the computing power center node, a workload energy consumption prediction model for the computing power center node, and an energy consumption prediction model for the cooling system. The three work together to establish a quantitative relationship between workload and energy consumption, realizing data sharing and interaction. In addition, by leveraging the powerful language generation and knowledge reasoning capabilities of generative artificial intelligence, flexible and efficient scheduling strategies are generated, significantly improving the flexibility and intelligence of the system and better adapting to complex and changing computing power demand scenarios.

[0077] As Figure 2 shown, Figure 2 a flowchart of a distributed computing power center load and energy consumption collaborative scheduling method based on artificial intelligence is disclosed.

[0078] Next, each step of the above solution will be introduced in detail in combination with Figure 2 :

[0079] In step S1, historical data of node workload, historical data of workload requirements of new computing power tasks, historical data of energy consumption of node workload, historical data of node environment, and historical data of energy consumption of the cooling system are collected and preprocessed.

[0080] In this step, historical data of node workload, historical data of workload requirements of new computing power tasks, historical data of energy consumption of node workload, historical data of node environment, and historical data of energy consumption of the cooling system are collected and preprocessed.

[0081] Exemplarily:

[0082] The historical data of node workload includes CPU usage rate, average load, memory occupancy, memory free space, memory usage rate, disk read and write speed, disk read and write times, disk I / O waiting time, number of running processes, task queue length, CPU time occupied by each process, number of transactions processed per second, amount of data transmitted per second, etc.

[0083] The historical data of workload requirements of new computing power tasks includes task type, task priority, data volume size, model complexity, number of users, frequency of user operations, number of instructions to be executed per second, number of floating-point operations per second, frequency of data reading and writing, task parallelism, etc.

[0084] The historical data of energy consumption of node workload includes overall node energy consumption within the input time step, average energy consumption of the node within the input time step, CPU energy consumption, memory energy consumption, storage device energy consumption, etc.

[0085] The node environmental historical data includes the environmental temperature of the node, the machine room temperature of the computing power center node, the difference between the weather temperature and the node machine room temperature, the environmental humidity, the machine room humidity, etc.

[0086] The energy consumption historical data of the cooling system includes the overall energy consumption of the node cooling system within the input time step, the average energy consumption of the node cooling system within the input time step, etc.

[0087] Furthermore, the preprocessing process includes data cleaning for abnormal data, numericalization of non-numerical data, and normalization of the data to remove noise and outliers, making the data have better consistency and comparability.

[0088] In step S2, based on the preprocessed historical data of the node workload and the historical data of the workload requirements of the new computing power tasks, the first training set and the first test set are divided, and a computing power center node workload prediction model is constructed using GAT and LSTM to obtain node workload prediction data.

[0089] This step includes: dividing the dataset based on the preprocessed historical data of the node workload and the historical data of the workload requirements of the new computing power tasks; constructing a computing power center node workload prediction model; training and optimizing the model.

[0090] Exemplarily, in this step, the dataset is randomly divided into a first training set and a first test set according to a ratio of 8:2.

[0091] As Figure 3 shown, the computing power center node workload prediction model includes two LSTM layers and two GAT layers; where:

[0092] The first LSTM layer is used to receive the historical data of the node workload of multiple historical time steps of multiple computing nodes, and transfer the features output by it to the first GAT layer to obtain the first feature through multiple attention heads;

[0093] The second LSTM layer is used to receive the historical data of the workload requirements of multiple new computing power tasks, and reshape the features output by it to obtain the second feature;

[0094] The second GAT layer is used to receive the combined result of the first feature and the second feature to obtain the historical data of the node workload of multiple historical time steps of multiple computing nodes.

[0095] Specifically:

[0096] The parameters involved in the model construction process are: the number of computing nodes Nodes_Num, Time_Step represents the output time step, Nodes_Features_1 represents the number of features of the load data of the computing nodes, and the batch size of the samples is Batch_Size.

[0097] The first LSTM layer: Its input is the load data of the Time_Step historical time steps of the original tasks of Nodes_Num computing nodes. With the powerful time series processing ability of the LSTM network, this layer can deeply mine and extract the time series features in the input historical load data. The number of hidden layer neurons in this layer is set to Nodes_Features_1, the input shape is [Nodes_Num×Time_Step, Nodes_Features_1], and the output structure is the features of the first LSTM layer with the shape of [Nodes_Num×Time_Step, Nodes_Features_1]. This processing method makes full use of the time correlation in the historical load data and provides strong support for subsequent analysis and prediction.

[0098] The first GAT layer: This layer is mainly responsible for automatically capturing the historical load dependency relationships between computing nodes. The number of attention mechanism heads in the first GAT layer is Attention_Head_1, and the number of hidden layer neurons is Hidden_Channels_1. Its input data has the shape of [Nodes_Num×Time_Step, Nodes_Features_1]. By setting the attention mechanism with Attention_Head_1 heads, this layer can deeply mine the features of the input data and convert it into an output with the shape of [Nodes_Num×Time_Step, Hidden_Channels_1×Attention_Head_1]. This design of multiple attention heads enables the model to capture data features from different angles, significantly improving the comprehensiveness and accuracy of feature extraction.

[0099] Second LSTM layer: This layer focuses on processing the load data of new computing power tasks. Assume that the maximum number limit of the load requirements of new computing power tasks is Task_Limit. The number of neurons in the hidden layer of this layer is Nodes_Features_1. This layer receives the load data of Task_Limit new tasks, and the task load feature is Nodes_Features_1. When the number of computing power tasks is less than Task_Limit, 0 is used to fill it up to Task_Limit. The input and output shapes of this layer are both [Task_Limit, Nodes_Features_1]. This process realizes the effective extraction of the features of new task load data, laying a foundation for the subsequent fusion with historical load data and further analysis.

[0100] Data structure reshaping and feature merging: To fuse the historical load data processed by the first GAT layer and the new task load data processed by the second LSTM layer, the output of the second LSTM layer needs to be reshaped into [Nodes_Num × Time_Step, Nodes_Features_1]. The data with the reshaped structure is merged with the structure of the output of the first GAT layer, which is [Nodes_Num × Time_Step, Hidden_Channels_1 × Attention_Head_1], in the sample dimension. The merged structure is [Nodes_Num × Time_Step, Hidden_Channels_1 × Attention_Head_1 + Nodes_Features_1].

[0101] Second GAT layer: The main function of this layer is to deeply fuse the historical load data processed by the first GAT layer and the new computing power task load data processed by the second LSTM layer and complement the dynamic load feature information. The number of attention mechanism heads in the second GAT layer is 1, and the number of neurons in the hidden layer is set to Nodes_Features_1. The input data structure of this layer is [Nodes_Num × Time_Step, Hidden_Channels_1 × Attention_Head_1 + Nodes_Features_1], and the output structure is [Nodes_Num × Time_Step, Nodes_Features_1], which is the data structure of the predicted workload data.

[0102] Furthermore, in this step, hyperparameter optimization of the workload prediction model for computing center nodes is also carried out through optimization algorithms such as grid search. After obtaining the optimal workload prediction model for computing center nodes, its output is saved for subsequent prediction of the workload of nodes.

[0103] Training process:

[0104] Before starting the formal training, determine the hyperparameters that need to be optimized by grid search, such as the learning rate, the number of attention mechanism heads, etc., and define the value ranges of these hyperparameters.

[0105] In each training epoch, set the model to training mode model_1.train(). For different hyperparameter combinations, perform the following operations:

[0106] Traverse the training data loader train_loader_1. For each batch of sample data, first clear the optimizer gradients optimizer.zero_grad().

[0107] Input the input data data_1.x and the calculation node correspondence data_1.edge into the computing power center node workload prediction model to obtain the output out_1, and calculate the loss loss by computing the measured load data data_1.y corresponding to the input data and out_1.

[0108] Backpropagate to calculate the gradients loss.backward(), and update the model parameters optimizer.step(). Accumulate the losses of each batch, and record the average loss of each training epoch under this hyperparameter combination.

[0109] Evaluation process:

[0110] After training all hyperparameter combinations, set the model to evaluation mode model_1.eval().

[0111] For each hyperparameter combination tested during the training process, traverse the test data loader test_loader_1. For each batch of data test_data_1, input the data into the model to obtain the output test_out_1, and calculate the loss test_loss.

[0112] Accumulate the losses of all test batches, and record the average loss of each hyperparameter combination on the test set. By comparing the average losses of the test set under different hyperparameter combinations, find the hyperparameter combination that optimizes the model performance, and use this to optimize the workload prediction model. Finally, print out the average loss of the test set under the optimal hyperparameter combination.

[0113] In step S3, based on the historical energy consumption data and the historical node environment data of the preprocessed node workload, combined with the node workload prediction data, divide the second training set and the second test set, and use GAT and LSTM to construct a computing power center node workload energy consumption prediction model to obtain the energy consumption prediction data of the node workload.

[0114] This step includes: dividing the dataset based on the historical energy consumption data of node workloads, the historical environmental data of nodes, and the predicted data of node workloads; constructing a prediction model for the energy consumption of node workloads in the computing power center; training and optimizing the model.

[0115] It should be noted that in this step, the predicted node load values of the prediction model for node workloads in the computing power center obtained in the previous step are combined with the historical load energy consumption and historical environmental data of the nodes and then input into the prediction model for node load energy consumption in the computing power center. This is mainly to improve the prediction accuracy of load energy consumption through multi-dimensional data fusion. The predicted node load values provide the dynamic trend of future loads, while the historical load energy consumption data reflects the energy consumption patterns of nodes under different loads, and the environmental data captures the impact of external conditions on energy consumption. Combining these data can more comprehensively analyze the complex relationship between workloads and load energy consumption, especially the direct impact of environmental factors on the heat dissipation efficiency and energy consumption of equipment.

[0116] Exemplarily, this step randomly divides the dataset into a second training set and a second test set in a ratio of 8:2. In addition, the model structure is constructed according to the format of the input data. The data structures of the second training set and the second test set are transformed according to the input format of the prediction model for workload energy consumption, and the prediction model for the energy consumption of node workloads in the computing power center is fitted and optimized until the mean absolute error of the loss function no longer decreases significantly.

[0117] As Figure 4 shown, the prediction model for the energy consumption of node workloads in the computing power center includes three LSTM layers and two GAT layers; where:

[0118] The third LSTM layer is used to receive the predicted data of node workloads at multiple historical time steps of multiple computing nodes to obtain the third feature;

[0119] The fourth LSTM layer is used to receive the historical energy consumption data of node workloads at multiple historical time steps of multiple computing nodes to obtain the fourth feature;

[0120] The fifth LSTM layer is used to receive the historical environmental data of nodes at multiple historical time steps, and after reshaping the data structure of the output features, obtain the fifth feature;

[0121] The third GAT layer is used to receive the combined result of the third feature, the fourth feature, and the fifth feature, and transfer it to the fourth GAT layer to obtain the predicted energy consumption data of node workloads at multiple historical time steps of multiple computing nodes.

[0122] Specifically:

[0123] The parameters involved in the model construction process are as follows: the number of computing nodes is Nodes_Num, Time_Step represents the output time step, Nodes_Features_1 represents the number of features of the predicted load data of the computing power center node workload prediction model, Nodes_Features_2 represents the number of features of the historical load energy consumption data of the computing nodes, Environment_Features_3 represents the number of features of the historical environment data of the computing nodes, and the batch size of the samples is Batch_Size.

[0124] The third LSTM layer: The number of neurons in the hidden layer of this layer is Nodes_Features_1. Its input is the Time_Step predicted load data of Nodes_Num computing nodes output from the computing power center node workload prediction model, and the input shape is [Nodes_Num×Time_Step, Nodes_Features_1]. This layer can deeply mine and extract the time series features in the input predicted load data, and the output structure is the features of [Nodes_Num×Time_Step, Nodes_Features_1].

[0125] The fourth LSTM layer: The number of neurons in the hidden layer of this layer is Nodes_Features_2. Its input is the Time_Step historical load energy consumption data of Nodes_Num computing nodes, and the input shape is [Nodes_Num×Time_Step, Nodes_Features_2]. This layer can deeply mine and extract the time series features in the input historical load energy consumption data, and the output structure is the features of [Nodes_Num×Time_Step, Nodes_Features_2].

[0126] The fifth LSTM layer: The number of neurons in the hidden layer of this layer is Environment_Features_3×Nodes_Num. Its input is the Time_Step historical environment data, and the input shape is [Time_Step, Environment_Features_3]. The output structure is the features of [Time_Step, Environment_Features_3×Nodes_Num].

[0127] Data structure reshaping: Reshape the output features of the fifth LSTM layer into [Nodes_Num×Time_Step, Environment_Features_3].

[0128] Feature merging: Merge the output features of the fifth LSTM layer after reshaping the structure with the outputs of the third LSTM layer with the structure of [Nodes_Num×Time_Step, Nodes_Features_1] and the fourth LSTM layer with the structure of [Nodes_Num×Time_Step, Nodes_Features_2] at the sample dimension. The merged structure is [Nodes_Num×Time_Step, Nodes_Features_1 + Nodes_Features_2 + Environment_Features_3].

[0129] The third GAT layer: The main function of this layer is to deeply integrate and complement the dynamic load feature information of the node prediction load data processed by the third LSTM layer, the node historical load energy consumption data processed by the fourth LSTM layer, and the historical environment data processed by the fifth LSTM layer. The number of attention mechanism heads in the third GAT layer is Attention_Head_3, and the number of neurons in the hidden layer is Hidden_Channels_3. Its input data is in the shape of [Nodes_Num×Time_Step, Nodes_Features_1 + Nodes_Features_2 + Environment_Features_3]. By setting the attention mechanism with Attention_Head_3 heads, this layer can perform in-depth feature mining on the input data and transform it into an output with the shape of [Nodes_Num×Time_Step, Hidden_Channels_3×Attention_Head_3].

[0130] The fourth GAT layer: The number of attention mechanism heads in the fourth GAT layer is 1, and the number of neurons in the hidden layer is set to Nodes_Features_2. The input data structure of this layer is [Nodes_Num×Time_Step, Hidden_Channels_3×Attention_Head_3], and the output structure is [Nodes_Num×Time_Step, Nodes_Features_2], which is the predicted load energy consumption data output by the computing power center node workload energy consumption prediction model.

[0131] Furthermore, in this step, hyperparameter optimization of the computing power center node workload energy consumption prediction model is also performed through optimization algorithms such as grid search. After obtaining the optimal computing power center node workload energy consumption prediction model, its output is saved for subsequent prediction of the workload energy consumption of the nodes.

[0132] Training process:

[0133] Before starting the formal training, determine the hyperparameters that need to be optimized through grid search, such as the learning rate, the number of attention mechanism heads, etc., and define the value ranges of these hyperparameters.

[0134] In each training epoch, set the model to the training mode model_2.train(). For different hyperparameter combinations, perform the following operations:

[0135] Traverse the training data loader train_loader_2. For each batch of sample data, first clear the optimizer gradients optimizer.zero_grad().

[0136] Input the input data data_2.x and the computing node correspondence data_2.edge into the computing power center node workload energy consumption prediction model to obtain the output out_2. Calculate the loss loss by computing the measured load data data_2.y corresponding to the input data and out_2.

[0137] Backpropagate to calculate the gradients loss.backward(), and update the model parameters optimizer.step(). Accumulate the losses of each batch, and record the average loss of each training epoch under this hyperparameter combination.

[0138] Evaluation process:

[0139] After training is completed for all hyperparameter combinations, set the model to the evaluation mode model_2.eval().

[0140] For each hyperparameter combination tested during the training process, traverse the test data loader test_loader_2. For each batch of data test_data_2, input the data into the model to obtain the output test_out_2, and calculate the loss test_loss.

[0141] Accumulate the losses of all test batches, and record the average loss of each hyperparameter combination on the test set. By comparing the average losses of the test sets under different hyperparameter combinations, find the hyperparameter combination that optimizes the model performance, and use this to optimize the computing power center node workload energy consumption prediction model. Finally, print out the average loss of the test set under the optimal hyperparameter combination.

[0142] In step S4, based on the preprocessed historical energy consumption data of the cooling system and the historical node environment data, combined with the energy consumption prediction data of the node workload, divide the third training set and the third test set, and use LSTM to construct an energy consumption prediction model for the cooling system.

[0143] This step includes: dividing the dataset based on the preprocessed historical energy consumption data of the cooling system, the historical node environment data, and the energy consumption prediction data of the node workload; constructing an energy consumption prediction model for the cooling system of the computing power center; training and optimizing the model.

[0144] It should be noted that there is usually a correlation between the load energy consumption and the cooling system energy consumption. An increase in the load energy consumption often leads to an increase in equipment heat generation, which in turn requires the cooling system to increase the cooling capacity, and the cooling system energy consumption also rises accordingly. Environmental data can assist in correcting the prediction of load energy consumption and cooling energy consumption, avoiding deviations caused by relying solely on historical cooling energy consumption and predicted load energy consumption. The double-layer LSTM layer can learn from historical and predicted data to uncover this causal relationship, providing a basis for subsequent system regulation.

[0145] Exemplarily, the dataset is randomly divided into a third training set and a third test set in a ratio of 8:2. In addition, according to the format of the input data for model construction, the data structures of the third training set and the third test set are transformed, and the time step of the input data is set; the energy consumption prediction model for the cooling system of the computing power center is fitted based on the transformed training set until the mean absolute error of the loss function no longer decreases significantly.

[0146] As Figure 5 shown, the LSTM for constructing the energy consumption prediction model of the cooling system includes two LSTMs; among them:

[0147] After reshaping the energy consumption prediction data of the node workload at multiple historical time steps of multiple computing nodes, the sixth feature is obtained; and the historical energy consumption data of the cooling system is used as the seventh feature, and the preprocessed historical node environment data is used as the eighth feature;

[0148] The sixth LSTM layer is used to receive the combined result of the sixth feature, the seventh feature, and the eighth feature, and transfer it to the seventh LSTM layer to obtain the energy consumption prediction data of the cooling system at multiple historical time steps.

[0149] Specifically:

[0150] Data structure reshaping: The predicted load energy consumption data of Time_Step for Nodes_Num computing nodes output from the node workload energy consumption prediction model of the computing power center, with the input shape shown as [Nodes_Num×Time_Step,Nodes_Features_2], is reshaped into [Time_Step,Nodes_Num×Nodes_Features_2].

[0151] Feature merging: Merge the predicted load energy consumption data after reshaping the structure, the historical cooling system energy consumption data with the structure [Time_Step, Cooling_Features_3], and the historical environment data with the structure [Time_Step, Environment_Features_3] at the sample dimension. The merged data structure is [Time_Step, Nodes_Num × Nodes_Features_2 + Cooling_Features_3 + Environment_Features_3].

[0152] The sixth LSTM layer: The role of this layer is to extract the coupling features of the predicted load energy consumption data, the historical cooling system energy consumption data, and the historical environment data. The number of neurons in the hidden layer of this layer is Nodes_Features_2 + Cooling_Features_3 + Environment_Features_3. Input the merged data from the previous step into the sixth LSTM layer. The input structure of this layer is [Time_Step, Nodes_Num × Nodes_Features_2 + Cooling_Features_3 + Environment_Features_3], and the output structure is [Time_Step, Nodes_Features_2 + Cooling_Features_3 + Environment_Features_3].

[0153] The seventh LSTM layer: This layer further mines deeper and more complex feature patterns and long-term dependence relationships in the data. The number of neurons in the hidden layer of this layer is Cooling_Features_3. Its input structure is [Time_Step, Nodes_Features_2 + Cooling_Features_3 + Environment_Features_3], and the output structure is [Time_Step, Cooling_Features_3], which is the data structure of the predicted cooling system energy consumption.

[0154] Furthermore, in this step, hyperparameter optimization of the computing power center cooling system energy consumption prediction model is also performed through optimization algorithms such as grid search. After obtaining the optimal computing power center cooling system energy consumption prediction model, save its output for subsequent prediction of the cooling system energy consumption of the prediction nodes.

[0155] Training process:

[0156] Before starting the formal training, determine the hyperparameters that need to be optimized by grid search, such as the learning rate, the number of neurons in the hidden layer, etc., and define the value ranges of these hyperparameters.

[0157] In each training epoch, set the model to training mode model_3.train(). For different combinations of hyperparameters, perform the following operations:

[0158] Iterate through the training data loader train_loader_3. For each batch of sample data, first clear the optimizer gradients optimizer.zero_grad().

[0159] Input the input data data_3.x and the computing node correspondence data_3.edge into the computing power center cooling system energy consumption prediction model to obtain the output out_3. Calculate the loss loss by computing the input data corresponding historical measured cooling system energy consumption data data_3.y and out_3.

[0160] Backpropagate to calculate the gradients loss.backward(), and update the model parameters optimizer.step(). Accumulate the losses for each batch, and record the average loss for each training epoch under this combination of hyperparameters.

[0161] Evaluation process:

[0162] After training is completed for all combinations of hyperparameters, set the model to evaluation mode model_3.eval().

[0163] For each combination of hyperparameters tested during the training process, iterate through the test data loader test_loader_3. For each batch of data test_data_3, input the data into the model to obtain the output test_out_3, and calculate the loss test_loss.

[0164] Accumulate the losses for all test batches, and record the average loss for each combination of hyperparameters on the test set. By comparing the average losses on the test set for different combinations of hyperparameters, find the combination of hyperparameters that optimizes the model performance, and use this to optimize the computing power center cooling system energy consumption prediction model. Finally, print the average loss on the test set under the optimal combination of hyperparameters.

[0165] In step S5, collect the workload demand data of the new computing power task, as well as the latest historical data of node workload, workload energy consumption, node environment, and the latest historical data of the cooling system energy consumption. Through the computing power center node workload prediction model, the computing power center node workload energy consumption prediction model, and the cooling system energy consumption prediction model, predict the node workload, workload energy consumption, and cooling system energy consumption data within the future time window.

[0166] In step S6, the workload within the future time window, the energy consumption of the workload, and the energy consumption data of the cooling system are filled into a preset prompt information template and used as the input for the generative artificial intelligence to obtain the final workload-energy consumption collaborative scheduling scheme for the distributed computing power center.

[0167] This step integrates the prediction results of the above three prediction models, namely the latest predicted data of the node workload, the latest predicted data of the energy consumption of the node workload, and the latest predicted data of the energy consumption of the node cooling system, to construct the prompt information. And the prompt information is input into the generative artificial intelligence. Exemplarily, the generative artificial intelligence here is GPT-4, to generate the workload-energy consumption collaborative scheduling scheme for the distributed computing power center, and the workload and energy consumption of the distributed computing power center are collaboratively scheduled according to this scheme.

[0168] Exemplarily, the construction of a feasible prompt information template is as follows:

[0169] According to the task requirements, a prompt information template for inputting into GPT-4 is constructed, leaving a data interface for the information to be filled. An example of the prompt information template is:

[0170] "The following are the relevant predicted data of the distributed computing power center. Please generate a workload-energy consumption collaborative scheduling scheme:

[0171] Latest predicted data of node workload: '{nodes_load}'

[0172] Latest predicted data of the energy consumption of the node workload: '{nodes_energy_consumption}'

[0173] Latest predicted data of the energy consumption of the cooling system:

[0174] '{cooling_system_energy_consumption}'

[0175] There is a new task in the current computing power center: '{nodes_calculation_task}'

[0176] Please give a specific workload-energy consumption collaborative scheduling scheme according to the above information, including which nodes the tasks are assigned to, whether the operating parameters of the nodes need to be adjusted, whether the operating strategy of the cooling system needs to be adjusted, etc."

[0177] Therefore, the construction process of the prompt information in this step is as follows:

[0178] Integrate the computing power demand data of the new computing power task, the latest prediction data of the node workload, the latest prediction data of the energy consumption of the node workload, and the latest prediction data of the energy consumption of the cooling system to construct a prompt message, and fill the integrated prediction data into the prompt message template. The corresponding filling relationship is as follows:

[0179] Computing power demand data of the new computing power task: nodes_calculation_task,

[0180] Latest prediction data of the node workload: nodes_load,

[0181] Latest prediction data of the energy consumption of the node workload: nodes_energy_consumption,

[0182] Latest prediction data of the energy consumption of the cooling system:

[0183] cooling_system_energy_consumption.

[0184] After determining the above prompt message, use it as the input of the generative artificial intelligence to obtain the final collaborative scheduling scheme for the workload and energy consumption of the distributed computing power center. Among them, the response of the generative artificial intelligence specifically includes: a function that encapsulates the OpenAI interface, and inputs the prompt message filled according to the prompt template into the generative artificial intelligence model to obtain the collaborative scheduling scheme for the workload and energy consumption of the distributed computing power center.

[0185] So far, the embodiment of the present invention has completed all the processes of the collaborative scheduling method for the load and energy consumption of the distributed computing power center based on artificial intelligence.

[0186] Embodiment 2:

[0187] As Figure 6 shown, the embodiment of the present invention provides a collaborative scheduling system for the load and energy consumption of a distributed computing power center based on artificial intelligence, including:

[0188] Data acquisition and preprocessing module, used to collect and preprocess the historical data of the node workload, the historical data of the workload demand of the new computing power task, the historical data of the energy consumption of the node workload, the historical data of the node environment, and the historical data of the energy consumption of the cooling system;

[0189] Model training module, used to divide the first training set and the first test set based on the preprocessed historical data of the node workload and the historical data of the workload demand of the new computing power task, and use GAT and LSTM to construct a computing power center node workload prediction model to obtain node workload prediction data;

[0190] Based on the historical energy consumption data of the node workload and the historical node environment data after preprocessing, combined with the predicted data of the node workload, divide the second training set and the second test set, and use GAT and LSTM to construct a prediction model for the energy consumption of the computing power center node workload to obtain the predicted energy consumption data of the node workload;

[0191] And based on the historical energy consumption data of the cooling system and the historical node environment data after preprocessing, combined with the predicted energy consumption data of the node workload, divide the third training set and the third test set, and use LSTM to construct a prediction model for the energy consumption of the cooling system;

[0192] The data prediction module is used to collect the workload demand data of the new computing power task, as well as the latest historical data of the node workload, the energy consumption of the workload, the node environment, and the latest historical data of the energy consumption of the cooling system. Through the computing power center node workload prediction model, the computing power center node workload energy consumption prediction model, and the cooling system energy consumption prediction model, predict the workload of the node, the energy consumption of the workload, and the energy consumption data of the cooling system within the future time window;

[0193] The solution generation module is used to fill the workload, the energy consumption of the workload, and the energy consumption data of the cooling system within the future time window into a preset prompt information template, and use it as the input of the generative artificial intelligence to obtain the final distributed computing power center workload-energy consumption collaborative scheduling solution.

[0194] Example 3:

[0195] An embodiment of the present invention provides a storage medium that stores a computer program for the collaborative scheduling of the load and energy consumption of a distributed computing power center based on artificial intelligence. Among them, the computer program enables a computer to execute the distributed computing power center load and energy consumption collaborative scheduling method as described in Example 1.

[0196] Example 4:

[0197] An embodiment of the present invention provides an electronic device, including:

[0198] One or more processors; a memory; and one or more programs, where the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The programs include those for executing the distributed computing power center load and energy consumption collaborative scheduling method as described in Example 1.

[0199] It is understandable that the distributed computing power center load and energy consumption collaborative scheduling system, storage medium, and electronic device provided by the embodiments of the present invention correspond to the distributed computing power center load and energy consumption collaborative scheduling method provided by the embodiments of the present invention. For the explanations, examples, beneficial effects, and other parts of the relevant content, reference can be made to the corresponding parts in the distributed computing power center load and energy consumption collaborative scheduling method, which will not be elaborated here.

[0200] In summary, compared with the prior art, the following beneficial effects are achieved:

[0201] 1. Improve the accuracy and adaptability of workload scheduling: Use GAT to model the nodes and their mutual relationships in the distributed computing power center as a graph structure, and capture the complex correlation relationships between nodes and the diversity of tasks through the graph attention mechanism. Combining the intelligent decision-making ability of generative artificial intelligence, it can dynamically adjust the scheduling strategy according to the real-time workload situation and task characteristics, effectively respond to the dynamic changes of the workload in a timely manner, and achieve precise workload balancing.

[0202] 2. Reduce the energy consumption management cost and break through the technical bottleneck: Through innovation at the software level, use GAT to construct an energy consumption prediction model to accurately predict the energy consumption of the computing tasks of nodes and the energy consumption of the cooling system. Based on these accurate energy consumption predictions, combined with the workload-energy consumption collaborative scheduling scheme generated by generative artificial intelligence, without relying on a large number of hardware upgrades, optimize the operating state and task allocation of nodes, reasonably control energy consumption, break through the technical bottleneck at the hardware level, and significantly reduce energy consumption at a low cost.

[0203] 3. Achieve effective coordination between workload and energy consumption: Through the collaborative work of the node workload prediction module, the energy consumption prediction module of the node workload, and the energy consumption prediction module of the cooling system, establish a quantitative relationship between workload and energy consumption, and achieve data sharing and interaction.

[0204] 4. Enhance flexibility and intelligence: Comprehensively consider various complex factors such as the network structure of the distributed computing power center, the computing power of node devices, and the impact of environmental factors on the workload. Utilize the powerful language generation and knowledge reasoning ability of generative artificial intelligence to generate flexible and efficient scheduling strategies, significantly improve the flexibility and intelligence of the system, and better adapt to complex and changing computing power demand scenarios.

[0205] 5. Eliminate the lag in energy management and achieve optimal overall energy consumption control: It can conduct a forward-looking analysis of changes in the workload, predict the trend of energy consumption changes in advance, and avoid the problem of extended task response time caused by lagged adjustment. The introduced generative artificial intelligence can quickly adapt to the dynamic changes in the workload and energy consumption of the computing power center. When facing sudden large-scale computing power tasks, it can quickly generate a reasonable scheduling plan to ensure the timely processing of tasks while controlling the growth of energy consumption.

[0206] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0207] The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for coordinated scheduling of load and energy consumption of distributed computing centers based on artificial intelligence, characterized in that: include: Collect and pre-process historical data on node workloads, historical data on workload requirements for new computing tasks, historical data on energy consumption of node workloads, historical data on node environments, and historical data on energy consumption of cooling systems; Based on the preprocessed historical data of the node workload and the historical data of the workload requirements of the new computing task, and dividing the first training set and the first test set, a computing center node workload prediction model is constructed using GAT and LSTM to obtain node workload prediction data; Based on the preprocessed energy consumption history data of the node workload and the node environment history data, combined with the node workload prediction data, and dividing the second training set and the second test set, GAT and LSTM are used to build a computing power center node workload energy consumption prediction model to obtain node workload energy consumption prediction data; Based on the preprocessed energy consumption history data of the cooling system and the node environment history data, combined with the energy consumption prediction data of the node workload, and dividing the third training set and the third test set, an energy consumption prediction model of the cooling system is constructed using LSTM; Collect workload demand data of new computing tasks, as well as the latest historical data of node workload, workload energy consumption, node environment, and cooling system energy consumption, and predict the node workload, workload energy consumption, and cooling system energy consumption data in the future time window through the computing center node workload prediction model, the computing center node workload energy consumption prediction model, and the cooling system energy consumption prediction model; The workload, energy consumption of the workload, and energy consumption data of the cooling system within the future time window are filled into a preset prompt information template and used as input for generative artificial intelligence to obtain the final distributed computing center workload-energy consumption collaborative scheduling plan.

2. The method for coordinated scheduling of load and energy consumption of a distributed computing center according to claim 1, characterized in that: The workload prediction model of the computing center node includes two LSTM layers and two GAT layers; wherein: The first LSTM layer is used to receive the node workload history data of multiple historical time steps of multiple computing nodes, and pass the output features thereof to the first GAT layer to obtain the first feature through multiple attention heads; The second LSTM layer is used to receive the historical data of workload requirements of multiple new computing tasks, and reshape the output features to obtain the second features; The second GAT layer is used to receive the combined result of the first feature and the second feature, and obtain node workload prediction data of multiple historical time steps of multiple computing nodes.

3. The method for coordinated scheduling of load and energy consumption of a distributed computing center according to claim 2, characterized in that: The computing center node workload energy consumption prediction model includes three LSTM layers and two GAT layers; wherein: The third LSTM layer is used to receive node workload prediction data of multiple historical time steps of multiple computing nodes to obtain a third feature; The fourth LSTM layer is used to receive energy consumption history data of node workloads of multiple historical time steps of multiple computing nodes to obtain a fourth feature; The fifth LSTM layer is used to receive the node environment historical data of multiple historical time steps, and reshape the output features to obtain the fifth feature; The third GAT layer is used to receive the combined result of the third feature, the fourth feature and the fifth feature, and pass it to the fourth GAT layer to obtain energy consumption prediction data of node workloads of multiple historical time steps of multiple computing nodes.

4. The method for coordinated scheduling of load and energy consumption of a distributed computing center according to claim 3, characterized in that: The LSTM constructed cooling system energy consumption prediction model includes two LSTMs; wherein: After reshaping the data structure of the energy consumption prediction data of the node workloads of multiple historical time steps of multiple computing nodes, the sixth feature is obtained; the energy consumption history data of the cooling system is used as the seventh feature, and the pre-processed node environment history data is used as the eighth feature; The sixth LSTM layer is used to receive the combined result of the sixth feature, the seventh feature and the eighth feature, and pass it to the seventh LSTM layer to obtain energy consumption prediction data of multiple historical time steps of the cooling system.

5. The method for coordinated scheduling of load and energy consumption of a distributed computing center according to claim 1, characterized in that: The generative artificial intelligence adopts GPT-4.

6. The method for coordinated scheduling of load and energy consumption of a distributed computing center according to any one of claims 1 to 5, characterized in that: The node workload history data includes any one of the following: CPU usage rate, average load, memory usage, memory free space, memory usage rate, disk read and write speed, disk read and write times, disk I / O waiting time, number of running processes, task queue length, CPU time occupied by each process, number of transactions processed per second, and amount of data transmitted per second, or any combination of several of them; and / or The historical data of workload requirements of the new computing task include any one of the task type, task priority, data size, model complexity, number of users, frequency of user operations, number of instructions to be executed per second, number of floating-point operations per second, frequency of data reading and writing, and task parallelism, or any combination of several of them; and / or The energy consumption history data of the node workload includes any one or a combination of any one of the overall energy consumption of the node within the input step, the average energy consumption of the node within the input step, the energy consumption of the CPU, the energy consumption of the memory, and the energy consumption of the storage device; and / or The node environment history data includes any one of the node's ambient temperature, the computing center node's computer room temperature, the difference between the weather temperature and the node's computer room temperature, the ambient humidity, and the computer room humidity, or any combination of several of them; and / or The energy consumption history data of the cooling system includes any one of the overall energy consumption of the node cooling system within the input step length and the average energy consumption of the node cooling system within the input step length, or a combination of any several of the above.

7. A distributed computing center load and energy consumption coordinated scheduling system based on artificial intelligence, characterized in that: include: The data collection and preprocessing module is used to collect and preprocess the historical data of node workload, the historical data of workload requirements of new computing tasks, the historical data of energy consumption of node workload, the historical data of node environment, and the historical data of energy consumption of cooling system; A model training module is used to construct a computing center node workload prediction model using GAT and LSTM based on the preprocessed node workload historical data and the historical data of the workload requirements of the new computing task, and to divide the first training set and the first test set to obtain node workload prediction data; Based on the preprocessed energy consumption history data of the node workload and the node environment history data, combined with the node workload prediction data, and dividing the second training set and the second test set, GAT and LSTM are used to build a computing power center node workload energy consumption prediction model to obtain node workload energy consumption prediction data; and for constructing an energy consumption prediction model for the cooling system using LSTM based on the preprocessed energy consumption history data of the cooling system and the node environment history data, combined with the energy consumption prediction data of the node workload, and dividing the third training set and the third test set; A data prediction module is used to collect workload demand data of new computing tasks, as well as the latest historical data of node workload, workload energy consumption, node environment, and cooling system energy consumption, and predict the node workload, workload energy consumption, and cooling system energy consumption data in the future time window through the computing center node workload prediction model, the computing center node workload energy consumption prediction model, and the cooling system energy consumption prediction model; The solution generation module is used to fill the workload, energy consumption of the workload, and energy consumption data of the cooling system in the future time window into a preset prompt information template, and use it as the input of the generative artificial intelligence to obtain the final distributed computing center workload-energy consumption collaborative scheduling solution.

8. A storage medium, characterized in that: It stores a computer program for coordinated scheduling of distributed computing power center load and energy consumption based on artificial intelligence, wherein the computer program enables the computer to execute the method for coordinated scheduling of distributed computing power center load and energy consumption as described in any one of claims 1 to 6.

9. An electronic device, characterized in that: include: one or more processors; Memory; And one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the programs include a method for executing the distributed computing center load and energy consumption coordinated scheduling method as described in any one of claims 1 to 6.

Citation Information

Cited By

  • Artificial intelligence processor control method for dynamic computing power scheduling

    CN121935026A