Machine learning based distributed system load and energy efficiency optimization method

CN122816888APending Publication Date: 2026-09-25XIAN DOUXIAOBAO MATERIALS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611026390.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-10
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有技术的不足,本发明提供了基于机器学习的分布式系统负载与能效优化方法,解决了现有任务调度机制未将计算节点空间热污染传导效应与滞后能耗变化纳入动作评估,导致分布式系统长期能效与负载分布无法达到最优平衡的问题

Benefits of technology

[0023]1、本发明通过集中调度服务器将中央处理器利用率数据、内存利用率数据和进风口温度数据重构为时空状态张量,结合安全掩码层将原始动作分布数据中属于散热受限计算节点范畴的概率数值强制归0,在模型输入层面完成了针对分布式计算环境的空间拓扑建模,并在动作输出端阻断了向超温节点分配新任务的数据链路,规避了任务调度引发局部计算节点温度过载的硬件运行风险。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122816888A_ABST
    Figure CN122816888A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of distributed computing, and discloses a distributed system load and energy efficiency optimization method based on machine learning, which comprises the following steps: reconstructing central processor utilization rate data, memory utilization rate data and air inlet temperature data into a space-time state tensor; a main strategy network generates original action distribution data, a safety mask layer combines the air inlet temperature data to execute 0 setting and normalization output safety action distribution data; state data is pushed into a delay queue to maintain a static state for storage; a reward attribution network combines dynamic network edge weight to output a contribution weight scalar, which is multiplied by a power increment to generate delay reward data; summing immediate reward data and delay reward data generates joint reward data to update a network model. The application blocks the data transmission path for distributing tasks to an over-temperature node through a safety mask layer, quantifies the heat transfer relationship between nodes by using a graph attention mechanism, and completes lagging refrigeration energy consumption punishment disassembly and distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of distributed computing technology, specifically to a method for optimizing the load and energy efficiency of distributed systems based on machine learning. Background Technology

[0002] In distributed computing systems, centralized scheduling servers need to allocate multi-dimensional tasks to different computing nodes. Existing task scheduling methods typically rely on CPU and memory utilization for task distribution, without performing spatial topology modeling for the distributed computing environment at the model input level. Existing scheduling networks lack a mandatory zeroing protection mechanism for computing nodes with limited thermal performance when outputting action probabilities. As the data link for continuously allocating new tasks to overheated nodes remains uninterrupted, the hardware risk of task scheduling causing local computing node overheating persists.

[0003] Within a distributed system, heat transfer exists between computing nodes. The heat pollution generated by tasks running on these nodes leads to an increase in the overall operating power of the cooling equipment. Existing scheduling reward mechanisms only calculate immediate latency data after task execution, lacking a delay queue for static data retention. The existing mechanisms do not utilize graph attention to quantify the heat transfer relationships between computing nodes, and cannot perform graph attention feature aggregation to output contribution weight scalars. The incremental cooling energy consumption generated over long-term operation cannot be non-linearly distributed to the target execution-related data group, making it difficult to track and assign responsibility for the delayed energy consumption caused by node heat pollution.

[0004] Because the delayed energy consumption caused by node thermal contamination is difficult to track and assign responsibility for, existing machine learning scheduling models only use short-term latency indicators from the task scheduling dimension for model iteration during the network connection parameter update phase. Existing scheduling methods cannot perform weighted summation operations on immediate latency indicators and delayed energy consumption penalty indicators, and cannot generate global joint reward data that integrates multi-dimensional indicators. The model iteration method that does not integrate long-term cooling energy consumption penalty indicators causes the network to tend to unilaterally shorten task processing time, resulting in a failure to achieve a multi-dimensional control balance between task processing time and overall operating power in the distributed system. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention provides a machine learning-based method for optimizing the load and energy efficiency of distributed systems. This method solves the problem that existing task scheduling mechanisms fail to incorporate the spatial thermal pollution transmission effect and delayed energy consumption changes of computing nodes into the action evaluation, resulting in the inability of distributed systems to achieve an optimal balance between long-term energy efficiency and load distribution.

[0006] To achieve the above objectives, this invention provides a machine learning-based method for optimizing the load and energy efficiency of distributed systems, comprising the following steps:

[0007] The centralized scheduling server reconstructs the central processing unit utilization data, memory utilization data, and air inlet temperature data into a spatiotemporal state tensor;

[0008] The main strategy network generates raw action distribution data based on the spatiotemporal state tensor and task feature vector. The security mask layer combines the air inlet temperature data to perform zeroing and renormalization operations to output safe action distribution data. The centralized scheduling server issues task allocation instructions.

[0009] Read the task completion time data and target time data to generate instant reward data, and merge the spatiotemporal state tensor, task feature vector, safety action distribution data and scheduling timestamp into an execution-related data group and push it into the delay queue cache along with the instant reward data.

[0010] The reward attribution network combines the edge weights of the dynamic network and the execution-related data group to perform graph attention feature aggregation operations to output a contribution weight scalar. The incremental running power data is multiplied by the contribution weight scalar to generate attribution delayed reward data.

[0011] A weighted summation operation is performed on the immediate reward data and the attribution-delayed reward data to generate global joint reward data. The main policy network reads the global joint reward data to update the connection parameters of its internal nodes.

[0012] In the machine learning-based distributed system load and energy efficiency optimization method provided in this invention, the multi-dimensional hardware state is reconstructed into a spatiotemporal state tensor, directly performing spatial topology modeling of the distributed computing environment at the input level. Combined with the forced zeroing protection mechanism of the security mask layer, the path for overheated nodes to accept new task assignments is directly blocked from the action output end, establishing hardware-level temperature red line protection. A delay queue is used to store task scheduling state data and is activated when the energy consumption lag effect becomes apparent. Combining a reward attribution network and dynamic network edge weights, the incremental cooling energy consumption generated over long-term operation is characterized and traced, completing the nonlinear joint optimization of real-time reward and delayed energy consumption penalty across multiple time scales.

[0013] To fully characterize the operating environment of the distributed system, the CPU utilization data, memory utilization data, and inlet temperature data are reconstructed into a spatiotemporal state tensor. This process includes: the out-of-band management controller acquiring CPU utilization data, memory utilization data, and inlet temperature data corresponding to the same discrete time step identifier and encapsulating them into a one-dimensional heterogeneous feature vector, which is then sent back to the centralized scheduling server; the centralized scheduling server aligns and arranges the CPU utilization data, memory utilization data, and inlet temperature data according to the discrete time step identifier, and generates a three-dimensional spatial grid by combining the spatial positional relationship defined by the three-dimensional Cartesian coordinate data, and merges and outputs a spatiotemporal state tensor containing the three-dimensional spatial topological relationship and the discrete time step identifier in the channel dimension.

[0014] During the task scheduling execution phase, the main policy network generates raw action distribution data based on the spatiotemporal state tensor and task feature vectors. This includes: a 3D convolutional neural layer independently performing convolution kernel extraction and pooling dimensionality reduction operations on the spatiotemporal state tensor to output the environment's implicit feature vector; a fully connected neural layer performing linear mapping operations on the task feature vector to output the task's implicit feature vector; and the main policy network performing concatenation operations on the environment's implicit feature vector and the task's implicit feature vector in the data dimension to generate a joint feature vector. The action mapping layer then performs matrix multiplication on the joint feature vector and calls a normalized exponential function for numerical processing to generate the raw action distribution data.

[0015] To avoid localized computing node temperature overload caused by task scheduling, the security mask layer combines inlet temperature data with zeroing and renormalization operations to output safe action distribution data. This includes: the security mask layer extracting pre-stored hardware limit temperature thresholds and pre-stored temperature safety margins, performing a subtraction operation to generate a safe temperature red line value, comparing the inlet temperature data with the safe temperature red line value; forcibly setting the probability values ​​of computing nodes with limited heat dissipation in the original action distribution data to zero, and performing a renormalization operation on the probability values ​​of computing nodes with safe heat dissipation, combined with extremely small normal numbers to prevent division by zero errors, to generate safe action distribution data.

[0016] In assessing the execution time of immediate scheduling, the process involves reading task completion time data and target execution time data to generate immediate reward data, and then pushing the execution-related data group and immediate reward data into the delay queue cache. This includes: the centralized scheduling server receiving the task completion confirmation signal and end timestamp data, performing a subtraction operation between the end timestamp data and the corresponding scheduling timestamp to generate task completion time data; the centralized scheduling server extracting the pre-stored basic reward constant and pre-stored timeout penalty coefficient, and calling the extreme value function and the exponential function with the natural constant as the base to generate immediate reward data based on the difference between the task completion time data and the target execution time data; when the scheduling timestamp bound to the immediate reward data is the same as the scheduling timestamp value extracted by traversing the delay queue, the centralized scheduling server pushes the immediate reward data into the delay queue and writes it into the reserved data segment of the target execution-related data group.

[0017] To ensure the synchronization and alignment of long-term energy efficiency data, after pushing the immediate reward data into the delay queue and writing it into the reserved data segment of the target execution associated data group, the following steps are also taken: the centralized scheduling server generates a status suspension signal and attaches a suspension status identifier to the target execution associated data group; when the delay queue detects the suspension status identifier, it locks the read and write permissions of the target execution associated data group, cuts off the corresponding data output link, and causes the corresponding data group to remain in a static retention state in the delay queue storage block until the settlement trigger instruction is received.

[0018] To quantify the heat transfer relationship between computing nodes, the steps for obtaining dynamic network edge weights include: a centralized scheduling server performing subtraction on the mean temperature values ​​corresponding to adjacent discrete time step identifiers to generate first-order difference data of the mean temperature; calculating the absolute value of the first-order difference data of the mean temperature and outputting the mean change rate; activating the reward attribution network when the cumulative number of consecutively read mean change rates less than the pre-stored steady-state judgment minimum values ​​is consistent with the pre-stored time node counting threshold; the reward attribution network performing data slicing on the inlet temperature data time series to separate the single-node temperature array; extracting a fixed-length data segment from the single-node temperature array according to the time window length parameter and calculating the mean value of the time dimension; and performing cross-correlation calculation on the data segments of the first computing node and the second computing node to output the dynamic network edge weights.

[0019] By combining graph attention mechanisms with network edge weights to decompose and allocate energy consumption penalties, the reward attribution network performs graph attention feature aggregation operations to output contribution weight scalars. This includes: the reward attribution network decomposes the safety action distribution data and spatiotemporal state tensors to generate node state feature vectors; concatenating the node state feature vectors of the first and second computation nodes in the data dimension; performing matrix multiplication on the concatenation result and the attention mapping weight matrix, and calling the LeakyReLU activation function for numerical processing; combining dynamic network edge weights to perform multiplication and summation calculations; and calling an exponential function with the natural constant as the base to perform division normalization to generate the thermal pollution conduction attention coefficient; performing weighted aggregation calculations on the thermal pollution conduction attention coefficient and the node state feature vectors to generate an aggregated state vector; and inputting the aggregated state vector into multilayer perceptron neurons to perform nonlinear mapping operations to output contribution weight scalars.

[0020] After energy consumption tracking is completed, the process of multiplying the incremental operating power data with the contribution weight scalar to generate attribution delay reward data includes: the centralized scheduling server performing a subtraction calculation on the current operating power data and the historical baseline operating power data to output the incremental operating power data; multiplying the incremental operating power data with the pre-stored energy consumption conversion constant to generate total lag energy consumption penalty data; multiplying the total lag energy consumption penalty data with the contribution weight scalar to generate attribution delay reward data; binding the attribution delay reward data with the corresponding scheduling timestamp to generate a settlement trigger instruction and sending it to the delay queue; locating the target execution associated data group retained internally in the delay queue; writing the attribution delay reward data into the supplementary data block; and completing the nonlinear decomposition and allocation of the total lag energy consumption penalty data to the target execution associated data group.

[0021] The network connection weights are updated by combining multidimensional reward data. A weighted summation operation is performed on the immediate reward data and the attribution delayed reward data to generate global joint reward data. The main policy network reads the global joint reward data to update the internal node connection parameters, including: the centralized scheduling server performs multiplication operations on the immediate reward data and the pre-stored first weight constant, and on the attribution delayed reward data and the pre-stored second weight constant, respectively. The output weighted values ​​of the immediate reward and delayed reward are then added to generate global joint reward data. The global joint reward data, scheduling timestamps, and the spatiotemporal state tensor, task feature vector, and safety action distribution data contained in the execution-related data group are serialized and recombined to generate single-session scheduling experience sample data, which are then merged and spliced ​​to generate a training batch dataset. The value evaluation network performs a forward mapping projection operation on the spatiotemporal state tensor and the task feature vector to output the estimated state action value. It then performs a chain-like differentiation operation on the policy loss function value and the value loss function value to generate the network parameter update gradient matrix, and performs backpropagation calculations to update the internal node connection parameters.

[0022] This invention provides a machine learning-based method for optimizing load and energy efficiency in distributed systems. It offers the following advantages:

[0023] 1. This invention reconstructs the central processing unit utilization data, memory utilization data, and air inlet temperature data into a spatiotemporal state tensor through a centralized scheduling server. Combined with a security mask layer, the probability values ​​of the original action distribution data belonging to the category of heat-constrained computing nodes are forcibly reduced to 0. At the model input level, spatial topology modeling for the distributed computing environment is completed, and at the action output end, the data link for allocating new tasks to overheated nodes is blocked, thus avoiding the hardware operation risk of local computing node temperature overload caused by task scheduling.

[0024] 2. This invention uses a delayed queue to keep the execution-related data group in a static state, and uses a reward attribution network combined with dynamic network edge weights to perform graph attention feature aggregation operations to output a contribution weight scalar. The incremental running power data is multiplied by the contribution weight scalar to generate attribution delayed reward data. The graph attention mechanism is used to quantify the heat transfer relationship between various computing nodes in the distributed system, and the nonlinear decomposition and allocation of the cooling energy consumption increment generated by long-term operation to the target execution-related data group is completed. This solves the problem of difficulty in tracking and assigning responsibility for lagging energy consumption caused by node thermal pollution conduction.

[0025] 3. This invention generates global joint reward data by performing a weighted summation operation on immediate reward data and attribution delayed reward data. The global joint reward data is then reorganized into single-time scheduling experience sample data and the internal node connection parameters of the main policy network are updated. During the network model parameter update stage, short-term time consumption assessment indicators of the task scheduling dimension and long-term cooling energy consumption penalty indicators of the environment dimension are integrated, which enables the distributed system to achieve a multi-dimensional control balance between task processing time and overall operating power. Attached Figure Description

[0026] Figure 1 This is a flowchart of the method of the present invention;

[0027] Figure 2 This is a topology diagram of the operating environment of the present invention;

[0028] Figure 3 This is a logic diagram of spatiotemporal feature extraction and security mask operation in this invention;

[0029] Figure 4 This is a logic diagram of the delayed queue retention and reward attribution calculation of the present invention;

[0030] Figure 5 This is a comparison chart of the changes in air inlet temperature data according to the present invention;

[0031] Figure 6This is a comparison chart of the changes in task completion time and operating power data for this invention. Detailed Implementation

[0032] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] See attached document Figure 1 and attached Figure 2 This invention provides a machine learning-based method for optimizing the load and energy efficiency of distributed systems. This method is implemented in an operating environment built upon a mature industrial hardware architecture. The operating environment includes a centralized scheduling server, network communication nodes, a cluster of computing nodes, and precision air conditioning equipment in the data center.

[0034] The compute node cluster contains multiple independently configured compute nodes. The compute node motherboard is equipped with an out-of-band management controller, and a temperature monitoring sensor is fixed on the surface of the compute node's air inlet. The out-of-band management controller is electrically connected to the compute node's central processing unit module and memory module, and collects central processing unit utilization data and memory utilization data, while the temperature monitoring sensor collects air inlet temperature data.

[0035] The out-of-band management controller, temperature monitoring sensor, and centralized scheduling server are connected to the network communication node, and the three of them form a data transmission link. The out-of-band management controller continuously transmits the central processing unit utilization data and memory utilization data to the centralized scheduling server in accordance with the baseboard management control protocol, and the temperature monitoring sensor simultaneously transmits the air inlet temperature data to the centralized scheduling server.

[0036] The precision air conditioning equipment in the computer room is internally encapsulated with an industrial control gateway. The industrial control gateway is connected to the network communication node. The centralized scheduling server reads the operating power data stored in the industrial control gateway through the network communication node, generates task allocation instructions according to the main strategy network, and sends them to the corresponding computing nodes. At the same time, it allocates internal storage space to receive and retain the operating status data of the computing node cluster and the operating power data of the precision air conditioning equipment in the computer room, and then calls the data in the storage space to perform optimization calculations.

[0037] The central scheduling server deploys a main policy network, a value assessment network, a security mask layer, a delay queue, and a reward attribution network. It receives task feature vectors from external sources and converts the CPU utilization data, memory utilization data, and air inlet temperature data collected in the storage space into spatiotemporal state tensors, which are then synchronously input into the main policy network. The main policy network performs feature extraction operations and outputs the raw action distribution data.

[0038] The raw action distribution data is transmitted directly from the main policy network to the security mask layer. The security mask layer synchronously reads the air inlet temperature data in the storage space, compares the air inlet temperature data with the set safety temperature red line value, performs a value reset and renormalization operation on the raw action distribution data based on the comparison result, and outputs the safety action distribution data. The centralized scheduling server extracts the safety action distribution data to generate task allocation instructions and distributes them to the network communication nodes.

[0039] The centralized scheduling server uses a delay queue to retain execution data. It pushes the execution-related data group, which is composed of spatiotemporal state tensor, task feature vector, safety action distribution data and scheduling timestamp, into the delay queue for caching. It continuously calculates the mean change rate of the air inlet temperature data time series. When the mean change rate is continuously less than the steady-state judgment minimum value, the reward attribution network is activated.

[0040] The reward attribution network extracts the execution-related data set cached in the delay queue, retrieves the inlet temperature data time series from the storage space, calculates the cross-covariance value, outputs the dynamic network edge weights, performs graph attention feature aggregation operation by combining the dynamic network edge weights and the execution-related data set, and outputs the contribution weight scalar. It centrally schedules the incremental operating power data of the precision air conditioning equipment in the server computer room, multiplies the incremental operating power data by the contribution weight scalar to obtain the attribution delay reward data, and feeds the attribution delay reward data back to the main policy network along the data transmission loop. The main policy network reads the attribution delay reward data and updates the internal node connection parameters.

[0041] See attached document Figure 1 The distributed system load and energy efficiency optimization method based on machine learning provided by this invention specifically includes the following steps:

[0042] The S100 centralized scheduling server receives CPU utilization data and memory utilization data transmitted by the out-of-band management controller through network communication nodes, and simultaneously receives air inlet temperature data continuously transmitted by temperature monitoring sensors. It then reconstructs the CPU utilization data, memory utilization data, and air inlet temperature data in the same time dimension into a three-dimensional multi-channel spatiotemporal state tensor.

[0043] The S200 and main strategy network synchronously read the spatiotemporal state tensor and the task feature vector input from external nodes. Based on the spatiotemporal state tensor and task feature vector, they perform feature extraction operations to generate raw action distribution data. The security mask layer reads the raw action distribution data and the air inlet temperature data, performs zeroing and renormalization operations, and outputs the security action distribution data. The centralized scheduling server extracts the security action distribution data and generates corresponding task allocation instructions, which are then sent to the computing node cluster through network communication nodes.

[0044] The S300 compute node cluster executes task allocation instructions and returns task completion time data. The centralized scheduling server reads the task completion time data and the preset target time data, calculates the difference between the task completion time data and the target time data, generates instant reward data, merges the spatiotemporal state tensor, task feature vector, safety action distribution data and scheduling timestamp into an execution-related data group, and pushes the execution-related data group and instant reward data into a delay queue for caching.

[0045] The S400 centralized scheduling server continuously calculates the mean change rate of the inlet temperature data time series. When the mean change rate is continuously less than the minimum value condition for steady state determination, the reward attribution network extracts the inlet temperature data time series in the delay queue, calculates the cross covariance value, outputs the dynamic network edge weights, retrieves the cached execution association data group in the delay queue, combines the dynamic network edge weights with the execution graph attention feature aggregation operation, outputs the contribution weight scalar, and the centralized scheduling server retrieves the incremental operating power data of the precision air conditioning equipment in the computer room, multiplies the incremental operating power data with the contribution weight scalar to generate attribution delay reward data.

[0046] The S500 centralized scheduling server extracts the immediate reward data and attribution delayed reward data at the same scheduling timestamp from the delay queue, performs a weighted summation operation on the immediate reward data and attribution delayed reward data to generate global joint reward data, and inputs the global joint reward data into the master policy network along the feedback loop. The master policy network reads the global joint reward data to calculate the gradient of the internal network loss function, and applies the gradient value to update the internal node connection parameters.

[0047] In this embodiment, for S100, the clock generator configured inside the centralized scheduling server continuously outputs discrete time step identifiers. The centralized scheduling server generates a status polling instruction based on the discrete time step identifiers and transmits it to the network communication node. The instruction is then distributed to all out-of-band management controllers included in the computing node cluster via the network communication node. After receiving the status polling instruction, the out-of-band management controller sends a read request to the central processing unit module through the baseboard management control communication link and obtains the central processing unit utilization data. Simultaneously, it sends a read request to the memory module and obtains the memory utilization data. At the same time, it detects the potential signal transmitted by the temperature monitoring sensor and converts it into air inlet temperature data.

[0048] The out-of-band management controller merges and encapsulates CPU utilization data, memory utilization data, and inlet temperature data corresponding to the same discrete time step into a one-dimensional heterogeneous feature vector. After appending the corresponding discrete time step identifier to the data header of the one-dimensional heterogeneous feature vector, it continuously transmits it back to the centralized scheduling server via network communication nodes. The centralized scheduling server receives the one-dimensional heterogeneous feature vector and performs data parsing operations. According to the discrete time step identifier, it aligns and arranges the CPU utilization data, memory utilization data, and inlet temperature data returned by different computing nodes in the time dimension and writes them into the storage space in categories.

[0049] The storage space is pre-configured with the three-dimensional Cartesian coordinate data of the corresponding computing node cluster. The centralized scheduling server reads the aligned and arranged central processing unit utilization data, memory utilization data and air inlet temperature data from the storage space, and synchronously reads the three-dimensional Cartesian coordinate data to generate a three-dimensional spatial mesh according to the spatial positional relationship defined by the three-dimensional Cartesian coordinate data.

[0050] The centralized scheduling server defines the three-dimensional spatial grid filled with central processing unit utilization data, memory utilization data, and air inlet temperature data as the first data channel, the second data channel, and the third data channel, respectively. It aligns and merges the data in the channel dimension and outputs a spatiotemporal state tensor containing three-dimensional spatial topological relationships and discrete time step identifiers. This tensor is then transmitted to the main policy network as the environmental baseline observation state for subsequent calculations.

[0051] For S200, refer to the appendix. Figure 3 The centralized scheduling server receives the externally issued task feature vector, which contains the instruction cycle value and the estimated resource requirement value, and transmits it to the main policy network synchronously with the spatiotemporal state tensor.

[0052] The three-dimensional convolutional neural layer and the fully connected neural layer configured inside the main policy network perform operations on the input spatiotemporal state tensor and task feature vector, respectively. The three-dimensional convolutional neural layer independently performs convolution kernel extraction and pooling dimensionality reduction operations and outputs the environment latent feature vector. The fully connected neural layer performs linear mapping operations and outputs the task latent feature vector.

[0053] The main policy network performs concatenation operations on the environment's implicit feature vector and the task's implicit feature vector in the data dimension to generate a joint feature vector. It then performs matrix multiplication on the joint feature vector through the internally configured action mapping layer, calls the normalization exponential function to perform numerical processing on the matrix multiplication results, and generates a set of probability values ​​to be distributed to each computing node in the computing node cluster to form the original action distribution data and output it.

[0054] The main policy network transmits the original action distribution data to the security mask layer. The security mask layer synchronously reads the inlet temperature data corresponding to the same discrete time step identifier from the storage space inside the centralized scheduling server. The centralized scheduling server has pre-stored the hardware limit temperature threshold and temperature safety margin. The security mask layer extracts the hardware limit temperature threshold and temperature safety margin and performs a subtraction operation to generate the safe temperature red line value.

[0055] The security mask layer extracts the air inlet temperature data of each computing node in the computing node cluster one by one and performs a numerical comparison operation with the safe temperature red line value. When the air inlet temperature data is greater than the safe temperature red line value, a heat dissipation limited flag is output, and when the air inlet temperature data is less than or equal to the safe temperature red line value, a heat dissipation safe flag is output.

[0056] In this embodiment, based on the already generated heat-limited and heat-safe identifiers, the security mask layer performs hard-constraint filtering on the original action distribution data. Specifically, the probability values ​​of nodes belonging to the heat-limited computing node category in the original action distribution data are forcibly set to 0; simultaneously, a re-normalization operation is performed on the probability values ​​belonging to the heat-safe computing node category. The normalization results from all grid nodes are then integrated to generate the safe action distribution data. The practical engineering significance of this forced zeroing operation lies in establishing a temperature redline barrier, preventing new tasks from being assigned to nodes on the verge of thermal runaway. The mapping formula included in the safe action distribution data is redefined as follows:

[0057] ;

[0058] In the formula, Assigning safety action distribution data to computing nodes The probability proportion; Assigning computation nodes to the original motion distribution data The probability proportion; For identifying discrete time steps; and These are the index numbers of the compute nodes within the compute node cluster; This represents the total number of compute nodes contained in the compute node cluster. For computing nodes Corresponding discrete time step identifier The inlet air temperature data; This is a pre-stored hardware limit temperature threshold; This is a pre-stored temperature safety margin; This is an indicator function that outputs 1 when the input condition is true and 0 when the condition is false. This represents a very small positive number to prevent division by zero errors, and its value range is set in

[10] . −6 10 −5Between ] . A minimal positive number is introduced to guard against an extreme condition: when a serious failure of the precision air conditioning equipment in the computer room causes the temperature at all air inlets of the computing node cluster to exceed the warning line, the indicator function... Output 0 values ​​across the entire grid. If minimal positive constants are lacking. As a fallback, a denominator approaching absolute zero will directly trigger a computational crash. Introducing extremely small positive numbers degenerates the probability distribution into a uniform distribution. Combined with the highest-level frequency reduction instructions from the centralized scheduling server, this ensures hardware safety under extreme conditions.

[0059] The centralized scheduling server extracts the security action distribution data output by the security mask layer and performs numerical sampling calculations. It selects the computing node with the highest probability value as the target node, encapsulates the externally issued task feature vector with the index number of the target node to generate a task allocation instruction, and routes the task allocation instruction to the corresponding computing node in the computing node cluster via the network communication node. The corresponding computing node receives the instruction and starts the computing load execution.

[0060] Within the parallel time node for generating task allocation instructions, the centralized scheduling server triggers data retention actions, collects the spatiotemporal state tensor and task feature vectors in the input main policy network, synchronously extracts the security action distribution data output by the security mask layer, and reads the scheduling timestamp corresponding to the current scheduling clock.

[0061] The centralized scheduling server merges and spatiotemporal state tensors, task feature vectors, safety action distribution data, and scheduling timestamps to generate execution-related data groups, and then transmits and pushes them into the delay queue. After receiving the execution-related data groups, the delay queue sorts and caches multiple execution-related data groups according to the time order of the scheduling timestamps, and keeps them until the reward settlement conditions are triggered.

[0062] For S300, after receiving the task allocation instruction, the computing node starts the operation. After completing the computing load matched by the task allocation instruction, it sends a task completion confirmation signal and end timestamp data to the network communication node. The network communication node sends the task completion confirmation signal and end timestamp data back to the centralized scheduling server. The centralized scheduling server receives the task completion confirmation signal and end timestamp data and retrieves the scheduling timestamp corresponding to the time when the task allocation instruction was generated.

[0063] The centralized scheduling server performs a subtraction operation between the end timestamp and the scheduling timestamp to generate task completion time data. Simultaneously, it extracts the target time data from within the task feature vector and performs difference calculations to generate instant reward data. The specific calculation formula for the instant reward data is defined as follows:

[0064] ;

[0065] In the formula, For the corresponding scheduling timestamp Real-time reward data; For the corresponding scheduling timestamp Task completion time data; For the corresponding scheduling timestamp Target time consumption data; This refers to the basic reward constant pre-existing in the centralized scheduling server; This refers to the timeout penalty coefficient pre-stored in the centralized scheduling server; For extrema functions, the extrema function is at and When the difference is less than or equal to 0, the constant output value is 0, and the extreme value function is at... and When the difference is greater than 0, output the numerical difference itself; It is an exponential function with the natural constant as its base.

[0066] The centralized scheduling server outputs real-time reward data, which is stored internally within the centralized scheduling server as a specific quantitative value of the degree of service level agreement satisfaction.

[0067] After the centralized scheduling server outputs the instant reward data, it extracts the scheduling timestamp bound to the instant reward data and sends a data matching instruction to the delay queue. The delay queue traverses the execution-related data groups in its internal cache according to the data matching instruction and extracts the included scheduling timestamps. The centralized scheduling server then compares the scheduling timestamp bound to the instant reward data with the extracted scheduling timestamps.

[0068] When the scheduling timestamp values ​​are the same, the centralized scheduling server locks the target execution associated data group in the delay queue, pushes the instant reward data into the delay queue and writes it into the reserved data segment of the target execution associated data group. After the data binding is completed, a status suspension signal is generated and a suspension status identifier is attached to the target execution associated data group.

[0069] When the delay queue detects a suspended status flag, it locks the read and write permissions of the target execution associated data group, cuts off the corresponding data output link, and causes the corresponding data group to remain in a static state within the delay queue storage block until the centralized scheduling server issues a settlement trigger instruction for the attribution delay reward data.

[0070] For S400, please refer to the appendix. Figure 4The centralized scheduling server continuously extracts the inlet temperature data corresponding to the discrete time step identifier from the storage space, merges the inlet temperature data into an inlet temperature data time series according to the time sequence, extracts the inlet temperature data time series, calculates the average temperature value of all calculation nodes under the same discrete time step identifier, performs a subtraction operation on the average temperature value corresponding to adjacent discrete time step identifiers to generate the first-order difference data of the average temperature, performs absolute value calculation on the first-order difference data of the average temperature, and outputs the rate of change of the average.

[0071] The centralized scheduling server has pre-stored the steady-state judgment minimum value and the time node count threshold. When the rate of change of the mean value of continuous reading is less than the steady-state judgment minimum value, the number of discrete time step identifiers that meet the conditions is accumulated. When the accumulated number is equal to the time node count threshold, the centralized scheduling server outputs a heat balance trigger command and sends it to the delay queue. The delay queue then releases the read and write permission lock status of the target execution associated data group.

[0072] The centralized scheduling server synchronously obtains the current operating power data and historical baseline operating power data of the precision air conditioning equipment in the computer room through network communication nodes. It performs a subtraction calculation on the current operating power data and the historical baseline operating power data, outputs the incremental operating power data, reads the pre-stored energy consumption conversion constant, performs a multiplication operation on the incremental operating power data and the energy consumption conversion constant, generates total lag energy consumption penalty data, and sends the heat balance trigger command to the reward attribution network. The reward attribution network receives the heat balance trigger command and enters the active operation state.

[0073] Once the reward attribution network is in an active state, it extracts the inlet temperature data time series from the centralized scheduling server, performs data slicing on the inlet temperature data time series according to the index number of the computing node, and separates the single-node temperature array corresponding to each computing node in the computing node cluster.

[0074] The reward attribution network is configured with a time window length parameter. Based on this parameter, a fixed-length data segment is extracted from the single-node temperature array. The mean value of the corresponding time dimension is calculated for each extracted data segment. The data segments and their time dimension mean values ​​corresponding to any first and second computational nodes are extracted. Cross-correlation is calculated for the data segments of the first and second computational nodes, and the dynamic network edge weights are output. The specific formula for calculating the dynamic network edge weights is defined as follows:

[0075] ;

[0076] In the formula, First computing node With the second computing node Dynamic network edge weights between them; For identifying discrete time steps; This is the identifier of the starting discrete time step corresponding to the data segment being extracted; The pre-configured time window length parameter; For computing nodes Corresponding discrete time step identifier The inlet air temperature data; For the second computing node Corresponding discrete time step identifier The inlet air temperature data; First computing node Extract the mean value of the time dimension corresponding to the data segment; For the second computing node Extract the mean value of the time dimension corresponding to the data segment.

[0077] The reward attribution network traverses all computing nodes in the cluster, performing multiplication and summation calculations on combinations of nodes to generate a complete dynamic network edge weight matrix. This dynamic network edge weight matrix is ​​stored in the internal neuron storage block of the reward attribution network for later retrieval.

[0078] The reward attribution network extracts the dynamic network edge weight matrix stored in the internal neuron storage block, and synchronously extracts the safety action distribution data and spatiotemporal state tensor corresponding to the same discrete time step identifier from the centralized scheduling server. The safety action distribution data and spatiotemporal state tensor are split and calculated according to the index number of the computing node to generate a node state feature vector that is independent for each computing node.

[0079] The reward attribution network extracts the node state feature vectors of any first computing node and the second computing node. In terms of data dimension, it performs concatenation calculation on the node state feature vectors of the first computing node and the second computing node. Internally, it pre-stores the attention mapping weight matrix and performs matrix multiplication operation on the result of the concatenation calculation and the attention mapping weight matrix.

[0080] The reward attribution network uses the LeakyReLU activation function to numerically process the result of matrix multiplication. It extracts the dynamic network edge weights corresponding to the first and second computational nodes within the dynamic network edge weight matrix. It then multiplies the value output by the LeakyReLU activation function with these dynamic network edge weights. Finally, it iterates through all computational nodes within the cluster, summing the results. A division normalization operation is then performed using an exponential function with the natural constant as the base, generating and outputting the thermal pollution conduction attention coefficient. The specific formula for calculating the thermal pollution conduction attention coefficient is defined as follows:

[0081] ;

[0082] In the formula, First computing node With the second computing node The thermal contamination conduction attention coefficient between them; It is an exponential function with the natural constant as its base; Use the LeakyReLU activation function; The reward is a pre-stored attention mapping weight matrix within the attribution network; First computing node The corresponding node state feature vector and the second computation node The data concatenation result of the corresponding node state feature vectors; First computing node With the second computing node Dynamic network edge weights between them; This represents the total number of compute nodes contained in the compute node cluster. This refers to the index number of the compute node used for traversal operations within the compute node cluster. First computing node The corresponding node state feature vector and the third computation node The data concatenation result of the corresponding node state feature vectors; First computing node With the third computing node The dynamic network edge weights between them.

[0083] The reward attribution network extracts the thermal pollution conduction attention coefficient and node state feature vector, performs weighted aggregation calculation on the thermal pollution conduction attention coefficient and node state feature vector to generate an aggregated state vector, which is configured with multilayer perceptrons. The multilayer perceptrons contain nonlinear activation functions. The reward attribution network inputs the aggregated state vector into the multilayer perceptrons. The multilayer perceptrons perform nonlinear mapping operation on the aggregated state vector and outputs the contribution weight scalar corresponding to each scheduling timestamp.

[0084] The reward attribution network transmits the contribution weight scalar to the centralized scheduling server. The centralized scheduling server receives the contribution weight scalar, synchronously extracts the total delayed energy consumption penalty data retained internally, performs a multiplication operation on the total delayed energy consumption penalty data and the contribution weight scalar to generate attribution delay reward data, binds the attribution delay reward data with the corresponding scheduling timestamp, and generates a settlement trigger instruction.

[0085] The centralized scheduling server sends the settlement trigger instruction to the delay queue. The delay queue receives the settlement trigger instruction, extracts the scheduling timestamp contained in the settlement trigger instruction, locates the internally retained target execution associated data group based on the extracted scheduling timestamp, extracts the attribution delay reward data, writes the attribution delay reward data into the reserved supplementary data block inside the target execution associated data group, completes the nonlinear disassembly and allocation of the total delayed energy consumption penalty data to the target execution associated data group, removes the suspended status flag attached to the target execution associated data group, and unlocks the read and write permissions of the target execution associated data group.

[0086] For S500, after the delayed queue releases the read / write permission lock on the target execution associated data group, it sends the target execution associated data group back to the centralized scheduling server. The centralized scheduling server receives and parses the target execution associated data group, extracts the immediate reward data from the reserved data segment of the target execution associated data group, and synchronously extracts the attribution delayed reward data from the supplementary data block of the target execution associated data group. By extracting various parameters within the same target execution associated data group, the immediate reward data and the attribution delayed reward data are aligned in the scheduling timestamp dimension.

[0087] The centralized scheduling server internally stores a first weighting constant and a second weighting constant. It performs a multiplication operation on the immediate reward data and the first weighting constant, outputting a weighted value for the immediate reward. It then performs a multiplication operation on the attribution delayed reward data and the second weighting constant, outputting a weighted value for the delayed reward. Finally, it performs an addition operation on the weighted values ​​for the immediate and delayed rewards to generate global joint reward data. The specific calculation formula for the global joint reward data is defined as follows:

[0088] ;

[0089] In the formula, For the corresponding scheduling timestamp Global joint reward data; For the corresponding scheduling timestamp Real-time reward data; For the corresponding scheduling timestamp Attribution of delayed reward data; The first weight constant pre-stored for the centralized scheduling server; The second weight constant is pre-stored for the centralized scheduling server.

[0090] The centralized scheduling server performs data serialization and recombination operations on the global joint reward data, scheduling timestamps, and spatiotemporal state tensors, task feature vectors, and safety action distribution data contained in the target execution association data group to generate single scheduling experience sample data. The single scheduling experience sample data is written into the internal experience playback storage pool block and is in a retention state waiting for the neural network to read and perform gradient update operations.

[0091] The centralized scheduling server continuously monitors the amount of single scheduling experience sample data retained within the experience replay storage pool block. When the data amount equals the preset batch threshold, it triggers random sampling and reading operations within the experience replay storage pool block to extract multiple single scheduling experience sample data and merge and splice them to generate a training batch dataset.

[0092] The centralized scheduling server performs data unpacking and splitting operations on the training batch dataset, separating mutually mapped and aligned spatiotemporal state tensors, task feature vectors, safety action distribution data, and global joint reward data from the training batch dataset. Internally configured with a main policy network and a value evaluation network, the separated spatiotemporal state tensors and task feature vectors are synchronously input into the value evaluation network. The value evaluation network performs forward mapping projection operations on the spatiotemporal state tensors and task feature vectors, and outputs the corresponding estimated state action value values.

[0093] The centralized scheduling server performs difference calculation on the estimated state action value and the global joint reward data, performs a square operation on the difference calculation result, generates a value loss function value, and re-inputs the spatiotemporal state tensor and task feature vector into the main policy network. The main policy network performs feature extraction and outputs new action probability distribution data. The centralized scheduling server combines the new action probability distribution data and the estimated state action value value to perform gradient ascent calculation and generate a policy loss function value.

[0094] The centralized scheduling server calls the internally integrated automatic differential optimizer. The automatic differential optimizer reads the values ​​of the value loss function and the policy loss function, performs chain-like differentiation on the values ​​of the value loss function and the policy loss function, generates the network parameter update gradient matrix, extracts the network parameter update gradient matrix, and performs backpropagation calculation along the neuron connection links between the main policy network and the value evaluation network.

[0095] The automatic differential optimizer performs numerical superposition and assignment overwriting on the weights of neurons in the main policy network and the value evaluation network. The centralized scheduling server completes the iterative update of the node parameters in the main policy network and the value evaluation network. After the parameter iterative update, the main policy network enters a standby running state, waiting to receive the spatiotemporal state tensor corresponding to the next discrete time step identifier and start a new round of distribution prediction calculation. The centralized scheduling server completes the machine learning parameter optimization calculation loop by retaining the experience sample data of a single scheduling and calculating the gradient matrix.

[0096] Specific application implementation:

[0097] Considering the high-concurrency task scheduling and management of job nodes in a specific implementation scenario, the constants and initial parameters set for the machine learning-based distributed system load and energy efficiency optimization method are as follows:

[0098] The hardware limit temperature threshold is set to 85.0℃; the temperature safety margin is set to 5.0℃; the minimum normal number to prevent division by zero errors is set to 0.00001; the basic reward constant is set to 10.0; and the timeout penalty coefficient is set to 0.5.

[0099] During the thermal offset calculation performed by the centralized scheduling server, high-risk probability values ​​in the original action distribution data are hard-constrained and removed. At a certain discrete time step, the centralized scheduling server internally extracts the allocation probability ratio corresponding to a certain first computing node as 0.45. The measured air inlet temperature data corresponding to the first computing node is extracted from the synchronous self-storage space and found to be 82.0℃. The centralized scheduling server subtracts the hardware limit temperature threshold from the temperature safety margin to obtain a safe temperature red line value of 80.0℃. Since the measured air inlet temperature data of 82.0℃ is greater than the safe temperature red line value of 80.0℃, the indicator function outputs a value of 0. The centralized scheduling server then substitutes the safe action distribution data mapping formula to calculate the updated probability ratio of the first computing node:

[0100] The centralized scheduling server uses the calculated 0 as a safety probability component after forcibly setting it to zero, and appends it to the safety action distribution data. This eliminates the risk of downtime caused by dispatching large-scale tasks to nodes on the verge of thermal runaway, and prevents further heat accumulation.

[0101] After the task allocation instruction is issued and executed, the centralized scheduling server dynamically quantifies the computational latency of a single computing node. The centralized scheduling server extracts the task completion time data returned by the network communication node as 102.0 ms. The target time data extracted from the synchronized task feature vector is 100.0 ms. The centralized scheduling server then substitutes the data into the instant reward formula to calculate the instant reward value corresponding to the current scheduling timestamp:

[0102] ;

[0103] The calculated instant reward value of 3.68 is called by the delay queue and written into the reserved data segment of the corresponding target execution associated data group. This is used to constrain the subsequent gradient update direction of the main policy network and ensure that the task allocation logic tends to favor low-latency nodes.

[0104] As multiple rounds of task allocation instructions are continuously issued and executed, the connection parameters of the internal nodes of the main policy network undergo multiple iterative updates. After the latest scheduling cycle is completed, the centralized scheduling server calculates the latest global joint reward data as 8.5. As the calculated latest global joint reward data gradually approaches an extreme value, the centralized scheduling server ends the random sampling read operation of the experience replay storage pool block and locks the current internal node connection parameters of the main policy network.

[0105] Actual operation tests were conducted and data were compared. Specific verification results are detailed in the appendix. Figure 5 and attached Figure 6 The experiment compared the continuous operation data of the traditional static heuristic scheduling method and the machine learning-based distributed system load and energy efficiency optimization method of this invention under the same computing node cluster and the same initial distribution of computing load. The experimental data is recorded in Table 1.

[0106] Table 1. Comparison of runtime data between traditional static heuristic scheduling methods and the method of this invention.

[0107] Method type Total task processing time (h) Maximum air inlet temperature (°C) Average daily power consumption (kWh) of precision air conditioning equipment in the computer room Number of times the heat dissipation limitation indicator is triggered. Traditional static heuristic scheduling method 45.5 84.6 350.5 125 Optimization method of the present invention 24.2 79.2 210.2 0

[0108] See attached document Figure 5 And Table 1, Appendix Figure 5 This diagram compares the convergence of inlet temperature data at various time points during the high-concurrency task scheduling process. The horizontal axis represents the discrete time step identifier, dimensionless, with a value range of 0 to 400. The vertical axis represents the inlet temperature data, in °C, with a value range of 70 to 90. The diagram includes inlet temperature data curves corresponding to the traditional control method and the method of this invention.

[0109] During the concurrency surge phase from discrete time step marker 150 to 250, the traditional static heuristic scheduling method did not perform hard constraint filtering on the inlet temperature data, resulting in a heat blind spot in target allocation. The inlet temperature data rapidly increased and exceeded the safe temperature threshold, reaching a maximum of 84.6℃, approaching the hardware's extreme temperature limit. The optimized method of this invention extracts the inlet temperature data through a safety mask layer and performs a re-normalization operation to generate safe action distribution data for allocation constraints. The corresponding inlet temperature data curve of this invention exhibits a truncated, gentle fluctuation trend, immediately falling back near discrete time step marker 200. The inlet temperature data is strictly controlled below the safe temperature threshold of 80.0℃, eliminating the localized downtime caused by continuous heat accumulation.

[0110] See attached document Figure 6 And Table 1, Appendix Figure 6 This diagram displays the operating power data and task completion time control of the precision air conditioning equipment in the underlying data center during the scheduling process. The horizontal axis represents the scheduling timestamp, dimensionless, ranging from 0 to 600. The left vertical axis represents the task completion time data, in milliseconds (ms), ranging from 30 to 120; the right vertical axis represents the operating power data, in watts (W), ranging from 150 to 300. The diagram includes a line graph of operating power data from the traditional method, a line graph of operating power data from this invention, and a curve showing the task completion time data from this invention.

[0111] Around 300, the total energy consumption of the computing node cluster surged, increasing the cooling load on the precision air conditioning equipment in the data center. Traditional static heuristic scheduling methods, lacking multi-dimensional environmental state awareness and reward delay attribution mechanisms, cause air conditioning commands to passively follow the heat load, resulting in prolonged high-power operation. The highest power displayed by the power data curve of the traditional method reached 295.0W and remained high.

[0112] The optimization method of this invention extracts the first-order difference data of the mean temperature through a reward attribution network, dynamically updates the edge weights of the dynamic network, and constructs a global joint reward data execution parameter iteration in the experience replay storage pool block. As can be seen from the power consumption data curve and the task completion time data curve of this invention, during periods of fluctuating task completion time, the centralized scheduling server automatically reduces the probability of tasks assigned to high-heat-dissipation nodes, strictly limiting the daily power consumption of the precision air conditioning equipment in the data center to 210.2 kWh. The power consumption data of this invention rapidly converges to the baseline in the latter half of the entire scheduling surge cycle, simultaneously ensuring the efficient execution of task allocation instructions and the energy efficiency optimization of the hardware environment.

Claims

1. A method for optimizing load and energy efficiency in distributed systems based on machine learning, characterized in that, Includes the following steps: The centralized scheduling server reconstructs the central processing unit utilization data, memory utilization data, and air inlet temperature data into a spatiotemporal state tensor; The main strategy network generates raw action distribution data based on the spatiotemporal state tensor and task feature vector. The security mask layer combines the air inlet temperature data to perform zeroing and renormalization operations to output safe action distribution data. The centralized scheduling server issues task allocation instructions. Read the task completion time data and target time data to generate instant reward data, and merge the spatiotemporal state tensor, the task feature vector, the safety action distribution data and the scheduling timestamp into an execution association data group and push it into the delay queue cache along with the instant reward data; The reward attribution network combines the dynamic network edge weights and the execution-related data group to perform graph attention feature aggregation operation to output a contribution weight scalar. The incremental running power data is multiplied by the contribution weight scalar to generate attribution delay reward data. A weighted summation operation is performed on the immediate reward data and the attribution delayed reward data to generate global joint reward data. The master policy network reads the global joint reward data to update the internal node connection parameters.

2. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, The CPU utilization data, memory utilization data, and inlet air temperature data are reconstructed into a spatiotemporal state tensor, including: The out-of-band management controller acquires the CPU utilization data, memory utilization data and air inlet temperature data corresponding to the same discrete time step identifier, encapsulates them into a one-dimensional heterogeneous feature vector, and sends them back to the centralized scheduling server. The centralized scheduling server aligns and arranges the CPU utilization data, memory utilization data, and air inlet temperature data according to the discrete time step identifier, generates a three-dimensional spatial grid by combining the spatial positional relationship defined by the three-dimensional Cartesian coordinate data, and merges and outputs the spatiotemporal state tensor containing the three-dimensional spatial topology relationship and the discrete time step identifier in the channel dimension.

3. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, The main policy network generates raw action distribution data based on the spatiotemporal state tensor and the task feature vector, including: The three-dimensional convolutional neural layer independently performs convolution kernel extraction and pooling dimensionality reduction operations on the spatiotemporal state tensor to output the environmental latent feature vector, and the fully connected neural layer performs linear mapping operation on the task feature vector to output the task latent feature vector. The main policy network performs a concatenation operation on the environment latent feature vector and the task latent feature vector in the data dimension to generate a joint feature vector. The action mapping layer performs matrix multiplication on the joint feature vector and calls the normalization exponential function for numerical processing to generate the original action distribution data.

4. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, The security mask layer, combined with the air inlet temperature data, performs a zeroing and renormalization operation to output safety action distribution data, including: The security mask layer extracts the pre-stored hardware limit temperature threshold and performs a subtraction operation with the pre-stored temperature safety margin to generate a safety temperature red line value, and performs a numerical comparison operation between the air inlet temperature data and the safety temperature red line value. The probability values ​​of the original action distribution data that belong to the category of heat-limited computing nodes are forcibly set to 0. The probability values ​​that belong to the category of heat-safe computing nodes are combined with the smallest normal numbers to prevent division by zero errors and a re-normalization operation is performed to generate the safe action distribution data.

5. The distributed system load and energy efficiency optimization method based on machine learning according to claim 1, characterized in that, Reading task completion time data and target time data to generate instant reward data, and pushing the execution-related data group and the instant reward data into the delay queue cache includes: The centralized scheduling server receives a task completion confirmation signal and an end timestamp data, and performs a subtraction operation between the end timestamp data and the corresponding scheduling timestamp to generate the task completion time data. The centralized scheduling server extracts the pre-stored basic reward constant and the pre-stored timeout penalty coefficient, and calls the extreme value function and the exponential function with the natural constant as the base to generate the instant reward data based on the difference between the task completion time data and the target time data. When the scheduling timestamp bound to the instant reward data is the same as the scheduling timestamp value extracted by traversing the delay queue, the centralized scheduling server pushes the instant reward data into the delay queue and writes it into the reserved data segment of the target execution associated data group.

6. The distributed system load and energy efficiency optimization method based on machine learning according to claim 5, characterized in that, After pushing the instant reward data into the delayed queue and writing it into the reserved data segment of the target execution associated data group, the process also includes: The centralized scheduling server generates a status suspension signal and attaches a suspension status identifier to the target execution associated data group; When the delay queue detects the suspended state identifier, it locks the read and write permissions of the target execution associated data group, cuts off the corresponding data output link, and causes the corresponding data group to remain in a static state in the delay queue storage block until a settlement trigger instruction is received.

7. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, The steps for obtaining the edge weights of the dynamic network include: The centralized scheduling server performs a subtraction operation on the mean temperature values ​​corresponding to adjacent discrete time step identifiers to generate first-order difference data of the mean temperature, and performs absolute value calculation on the first-order difference data of the mean temperature to output the mean change rate. The reward attribution network is activated when the cumulative number of consecutively read mean change rates less than the pre-stored steady-state minimum values ​​is equal to the pre-stored time node counting threshold. The reward attribution network performs data slicing on the inlet temperature data time series to separate single-node temperature arrays. Based on the time window length parameter, a fixed-length data segment is extracted from the single-node temperature array, and the mean value of the time dimension is calculated. The cross-correlation is calculated for the data segments of the first computing node and the second computing node, and the dynamic network edge weights are output.

8. The distributed system load and energy efficiency optimization method based on machine learning according to claim 1, characterized in that, The reward attribution network performs graph attention feature aggregation operations and outputs contribution weight scalars, including: The reward attribution network splits the safety action distribution data and the spatiotemporal state tensor to generate node state feature vectors, and performs concatenation calculation on the node state feature vectors of the first computing node and the node state feature vectors of the second computing node in the data dimension. The splicing calculation result and the attention mapping weight matrix are subjected to matrix multiplication and the LeakyReLU activation function is called for numerical processing. The dynamic network edge weights are combined with multiplication and summation calculations. The exponential function with the natural constant as the base is called to perform division normalization to generate the thermal pollution conduction attention coefficient. The thermal pollution conduction attention coefficient and the node state feature vector are weighted and aggregated to generate an aggregated state vector. The aggregated state vector is then input into a multilayer perceptron to perform a nonlinear mapping operation and output the contribution weight scalar.

9. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, The attribution delay reward data is generated by multiplying the incremental operating power data by the scalar contribution weight, including: The centralized scheduling server performs a subtraction calculation between the current operating power data and the historical baseline operating power data to output the incremental operating power data. It then performs a multiplication operation between the incremental operating power data and the pre-stored energy consumption conversion constant to generate total lag energy consumption penalty data. The total delayed energy consumption penalty data and the contribution weight scalar are multiplied to generate the attribution delay reward data. The attribution delay reward data is then bound to the corresponding scheduling timestamp to generate a settlement trigger instruction, which is then sent to the delay queue. The delay queue locates the target execution associated data group retained inside, writes the attribution delay reward data into the supplementary data block, and completes the nonlinear decomposition and allocation of the total lag energy consumption penalty data to the target execution associated data group.

10. The method for optimizing load and energy efficiency in a distributed system based on machine learning according to claim 1, characterized in that, A weighted summation operation is performed on the immediate reward data and the attribution-delayed reward data to generate global joint reward data. The master policy network reads the global joint reward data to update its internal node connection parameters, including: The centralized scheduling server performs multiplication operations on the instant reward data and the pre-stored first weight constant, and on the attribution delayed reward data and the pre-stored second weight constant, respectively, and performs addition operations on the output instant reward weighted value and delayed reward weighted value to generate the global joint reward data; The global joint reward data, the scheduling timestamp, the spatiotemporal state tensor contained in the execution association data group, the task feature vector, and the safety action distribution data are serialized and recombined to generate single scheduling experience sample data, which are then merged and spliced ​​to generate a training batch dataset. The value assessment network performs a forward mapping projection operation on the spatiotemporal state tensor and the task feature vector to output the estimated state action value. It then performs a chain-like differentiation operation on the policy loss function value and the value loss function value to generate network parameter update gradient matrix and performs backpropagation calculation to update the internal node connection parameters.